REVIEW 3 major objections 4 minor 13 references
High-z stellar masses can be recovered robustly with JWST photometry
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that the stellar masses of z = 5–10 galaxies can be recovered from JWST NIRCam photometry to within roughly a factor of three by a standard SED-fitting code, and that the residual mass-dependent biases are driven by poor…
desk verdict Solid controlled mock-recovery study; the recovery statistics hold, but the emission-line mechanism and the Narayanan contrast are weaker than the abstract implies. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a forward-modelling chain with known ground truth. SPHINX20 provides simulated galaxies with BPASS v2.2.1 stellar spectra, CLOUDY-based emission-line luminosities, and Rascas Monte-Carlo dust radiative transfer along ten lines of sight; those spectra are convolved with the eight PRIMER NIRCam filter curves to make noise-free synthetic photometry, which is then refitted with BAGPIPES using BC03 stellar templates, CLOUDY nebular emission, and a flexible Salim dust-attenuation model. The load-bearing diagnostic is the offset $\Delta M_\star = \log_{10}(M_{\star,\mathrm{fitted}}/M_{\star,\mathrm{true}})$ plotted against true mass, specific star-formation rate, and H$\alpha$/[OIII] equivalent width. The identified mechanism is the emission-line bias: at $z = 5\!-\!8$, H$\alpha$ and [OIII] enter the F277W, F356W, F410M and F444W bands, and when the fit under-produces these strong lines it substitutes an older, more massive stellar population whose redder continuum matches the line-boosted photometry.
What would settle it
Measure H$\alpha$ and [OIII] equivalent widths spectroscopically for a sample of $z = 5\!-\!8$ galaxies and compare them with the equivalent widths returned by broadband SED fits of the same objects. If the paper's mechanism is right, the galaxies with the largest true line equivalent widths should show the largest photometric mass overestimates relative to masses derived with the line fluxes pinned to their spectroscopic values; seeing no such correlation, or seeing the mass-dependent trends persist when the fits are forced to match the measured lines, would undercut the claim that poor line modelling is what drives the bias.
Extended reading notes
Core claim
The central claim, stated in the abstract and Section 4.1, is that stellar masses at $z = 5\!-\!10$ are recovered robustly by JWST photometry: for SPHINX20 galaxies with $M_\star \sim 10^7\!-\!10^9\,\mathrm{M}_\odot$, fitting the forward-modelled NIRCam photometry with BAGPIPES yields median offsets below about 0.4 dex for every star-formation-history parametrisation, and 84–100% of masses within 0.5 dex. This is in direct contrast to the recent claim that stellar masses at these redshifts can be underestimated by as much as an order of magnitude. The paper further claims that the residual biases are driven by a specific mechanism: strong nebular emission lines (H$\alpha$ and [OIII] at $z = 5\!-\!8$) fall inside the red NIRCam filters, and when the fitting code cannot reproduce their equivalent widths it compensates with an older, more massive stellar population with a higher mass-to-light ratio. The same bias works in reverse at the high-mass end, where the code slightly overestimates line strengths and returns younger, less massive populations. These systematic trends exist for all six star-formation-history parametrisations and tilt the inferred stellar mass function, undercounting massive galaxies (by up to about 1 dex at $z \le 7$) and overcounting low-mass ones (by up to about 0.5 dex at $z \ge 8$).
Load-bearing premise
The load-bearing premise is that the SPHINX20 mock galaxies, including their emission-line strengths, dust-star geometries, and the SFR $> 0.3\,\mathrm{M}_\odot\,\mathrm{yr}^{-1}$ selection that puts them in the sample, together with the deliberately idealised fitting setup (redshifts fixed at the true values, noise-free photometry with 10% uncertainties, no AGN) are representative enough of real JWST observations for the recovery statistics and the line-driven bias to carry over to observed galaxies.
Editorial extensions
If this is right
- JWST-only NIRCam data can support stellar mass measurements at $z = 5\!-\!10$ at the factor-of-three level, so the pessimistic reading of recent simulation-based work is not the whole story.
- Survey-level stellar mass functions derived from broadband fitting are tilted by the mass-dependent bias: massive galaxy number densities are undercounted by up to about 1 dex at $z \le 7$, while low-mass number densities are inflated at $z \ge 8$.
- The choice of star-formation-history parametrisation makes little difference at these redshifts, because all six models fit the photometry comparably and the Bayesian information criterion even favours the simplest single-burst prior.
- Adding the NIRCam medium bands substantially improves recovery, for example raising the fraction of $z = 8$ galaxies recovered within 0.5 dex from 76% to 91%, by anchoring the continuum in filters free of strong line contamination.
- Because overestimation tracks rising recent star formation and high specific star-formation rates, star-forming galaxies bear the brunt of the bias; if the effect persists toward $z \sim 4$ it would inflate the apparent quenched fraction.
Reading between the lines
- The mechanism named by the paper implies the bias is not specific to BAGPIPES: any code fitting broad-band photometry with standard nebular-emission templates faces the same degeneracy between strong lines and an old, massive stellar population, so the disagreement with the pessimistic study may owe more to the simulated galaxies or the fitting setup than to the code itself.
- If real $z = 5\!-\!8$ galaxies have even stronger line emission than SPHINX20 predicts, as some JWST spectroscopy suggests, the low-mass overestimates in observed samples could exceed the roughly 0.5 dex seen here, making spectroscopic line constraints a more efficient safeguard than deeper photometry.
- The tilt in the inferred stellar mass function runs opposite to Eddington-style scatter around a steep mass function, so the two effects partially cancel at the low-mass end; separating them requires modelling the full scatter distribution rather than the median offset.
- A direct, cheap extension would be to refit the same photometry with line equivalent widths fixed to the simulated truth: if the mass-dependent trends vanish, the emission-line mechanism is confirmed and filter sets or fitting priors could be optimised to suppress the bias.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper tests whether stellar masses of high-redshift galaxies can be recovered robustly from JWST NIRCam photometry. Using 1,013 SPHINX20 galaxies at z=5-10, the authors forward-model JWST PRIMER photometry from radiative-transfer spectra and fit it with BAGPIPES under six star-formation-history models: single burst, exponential, delayed exponential, double power law, continuity, and bursty continuity. They report that recovered stellar masses are generally within a factor of ~3 of the true values, with the large majority within 0.5 dex, and that these results contrast with the order-of-magnitude underestimates claimed by Narayanan et al. (2024). They identify mass-dependent biases: overestimation for M*<~10^8 Msun and underestimation for M*>~10^9 Msun. They attribute the low-mass overestimation to BAGPIPES fitting strong emission lines poorly and show that these biases tilt the inferred stellar mass function. An appendix demonstrates that adding more NIRCam medium bands improves recovery.
Significance. If the main result holds, the paper provides a useful best-case benchmark for high-z stellar mass recovery from JWST photometry and a concrete mechanistic explanation for mass-dependent biases that can affect stellar mass functions. The study has clear strengths: it uses an external simulation with known ground-truth masses, it tests six SFH parametrizations, it forward-models photometry with radiative transfer, and it makes quantitative recovery statistics easy to interpret. The medium-band comparison in Appendix A is a particularly useful practical result. The central caveats are that the test is deliberately idealized (fixed true redshift, noise-free photometry with assigned 10% uncertainties, no MIRI, no AGN, and SFR-selected sample), and, as discussed below, the emission-line mechanism is partly circular because the same photoionization code is used to generate the mock line fluxes and the BAGPIPES nebular emission.
major comments (3)
- [Sections 2.2-2.3 and Figure 5] The mechanistic claim that the mass-dependent bias arises because BAGPIPES 'poorly models the impact of strong emission lines' is not independently established, because the mock line luminosities are generated with CLOUDY-based models (Section 2.2, citing Choustikov et al. 2024) while BAGPIPES's nebular emission is also based on CLOUDY (Section 2.3, citing Byler et al. 2017). Figure 5 therefore demonstrates a mismatch between the fitted and true line equivalent widths, but that mismatch could be partly due to shared CLOUDY assumptions rather than a generic limitation of SED fitting. To make the mechanism load-bearing, the authors should repeat the EW comparison using an independent line-emission model (or a line-luminosity calibration from observed high-z galaxies) for the mock spectra, or at least state explicitly that the diagnosis is conditioned on the CLOUDY framework.
- [Abstract and Section 4.1] The abstract states that '>90% of masses are recovered to within 0.5 dex', but Section 4.1 reports that for the Bursty Continuity model the fractions are 99, 100, 98, 84, 100 and 100% at z=5-10, so the z=8 value is 84%. The quantitative headline is therefore internally inconsistent. The abstract should either state the per-redshift percentages or use a threshold (e.g., '>84%' or '>90% except at z=8') that is actually supported by the results.
- [Sections 3.1-3.3 and 4.2.2] The population-level recovery statistics and the SMF-tilting conclusion are obtained under a deliberately best-case setup: the redshift is fixed to the true value, the photometry is noise-free with 10% assigned uncertainties, and the sample is selected by SFR10>0.3 Msun/yr. The paper acknowledges these choices, but the abstract's unqualified claim that stellar masses 'can be recovered robustly' with JWST photometry goes beyond what the idealized test can support. A concrete test with photometric redshift errors and realistic noise (or at least an explicit statement that the result is an upper bound on recovery quality) is needed to make the observational summary accurate.
minor comments (4)
- [Throughout] The simulation name is written inconsistently as 'Sphinx20' in the text and 'SPHINX20' in many figure captions and headings; please use a single spelling.
- [Section 3.1] There is a typographical issue in the sentence beginning 'At (M★ ≲ 107 M⊙)', where the opening parenthesis appears unpaired.
- [Figure 6] The lower panels use 'log10(N)' without defining the base or the meaning of the vertical axis; please label the axis as log10(N_inferred/N_true) for clarity.
- [Section 2.3] The phrase 'SNR of 10 is achieved in each band' in Section 5 is equivalent to the 10% uncertainties used earlier, but the connection is not stated; make the equivalence explicit.
Circularity Check
No significant circularity: stellar-mass recovery is an external benchmark against simulation-intrinsic ground truth, and the shared-CLOUDY mechanism caveat is a generalizability concern, not a circular reduction.
full rationale
The central claim is tested by comparing BAGPIPES-fitted stellar masses with SPHINX20 intrinsic stellar masses. The ground truth is fixed by the simulation's star formation and stellar particle bookkeeping, independent of BAGPIPES, and the synthetic photometry is forward-modeled from dust-attenuated spectra (Section 2.2), so no fitted parameter is renamed as a prediction. The mass-dependent bias is diagnosed by correlating Delta M* with sSFR, SFR10/SFR100, and line equivalent widths (Section 3.2, Figures 4-5); the line-EW test does compare quantities generated with the same photoionization code family (mock lines from CLOUDY via Choustikov et al. 2024; BAGPIPES nebular emission also CLOUDY via Byler et al. 2017), but the demonstrated failure of BAGPIPES to match the line fluxes is a real within-framework mismatch, not a tautology. The paper explicitly notes the stellar templates differ from those in the simulation (Section 2.3), and Section 4.2.2 flags the need to test different stellar population synthesis templates and emission-line modelling methods, acknowledging the generalizability limit. The use of the SPHINX20 simulation (Katz et al. 2023, with an overlapping author) is a normal self-citation to public, externally reproducible data and does not force the recovery statistics. No equation or fitted parameter reduces to the target result; the only caveats are representativeness of the mock galaxies and the 'best-case' fitting setup, which the paper states explicitly. Therefore no circularity is present under the enumerated patterns.
Assumptions & free parameters
free parameters (6)
- Photometric uncertainty on noise-free fluxes =
10% (S/N=10)
- Continuity prior scale and degrees of freedom =
sigma=0.3, nu=2 (continuity); sigma=1, nu=2 (bursty continuity)
- Non-parametric SFH time bin edges =
[0, 10, 30, 100, 300, t_max] Myr
- Log stellar mass prior range =
[5.0, 12.0] log10(M_sun/M_sun), uniform
- Dust attenuation prior =
A_V=[0.0, 2.0], bump strength fixed to 0, birth cloud factor 2
- SPHINX20 sample selection threshold =
SFR10 > 0.3 M_sun/yr
assumptions (5)
- domain assumption SPHINX20 intrinsic galaxy properties (stellar masses, SFRs) are a valid ground truth for testing SED fitting at z=5-10.
- domain assumption The forward-modelled spectra (BPASS stellar SEDs, CLOUDY emission lines, Rascas dust radiative transfer with SMC dust) faithfully represent the SEDs JWST would observe, including emission-line strengths and dust-star geometry.
- domain assumption BAGPIPES with the stated priors (BC03 2016 templates, Kroupa IMF, CLOUDY nebular emission with ISM grains, Salim et al. dust parametrization) is a representative 'commonly-used SED fitting code' for JWST observers.
- ad hoc to paper Fixing the redshift at the true value and using noise-free photometry with assigned 10% uncertainties is an appropriate 'best-case' benchmark for isolating SFH effects.
- standard math The BIC (BIC = k log n - 2L) is an appropriate criterion for comparing SFH parametrisations in this setting.
Cite this review
Pith. "Pith review of High-z stellar masses can be recovered robustly with JWST photometry." pith.science (2026). https://pith.science/paper/IHB75F4J
@misc{pith2026241202622,
author = {Pith},
title = {Pith review of: High-z stellar masses can be recovered robustly with JWST photometry},
year = {2026},
howpublished = {\url{https://pith.science/paper/IHB75F4J}},
note = {Machine review of arXiv:2412.02622}
}
read the original abstract
Robust inference of galaxy stellar masses from photometry is crucial for constraints on galaxy assembly across cosmic time. Here, we test a commonly-used Spectral Energy Distribution (SED) fitting code, using simulated galaxies from the SPHINX20 cosmological radiation hydrodynamics simulation, with JWST NIRCam photometry forward-modelled with radiative transfer. Fitting the synthetic photometry with various star formation history models, we show that recovered stellar masses are, encouragingly, generally robust to within a factor of ~3 for galaxies in the range M*~10^7-10^9M_sol at z=5-10. These results are in stark contrast to recent work claiming that stellar masses can be underestimated by as much as an order of magnitude in these mass and redshift ranges. However, while >90% of masses are recovered to within 0.5dex, there are notable systematic trends, with stellar masses typically overestimated for low-mass galaxies (M*<~10^8M_sol) and slightly underestimated for high-mass galaxies (M*>~10^9M_sol). We demonstrate that these trends arise due to the SED fitting code poorly modelling the impact of strong emission lines on broadband photometry. These systematic trends, which exist for all star formation history parametrisations tested, have a tilting effect on the inferred stellar mass function, with number densities of massive galaxies underestimated (particularly at the lowest redshifts studied) and number densities of lower-mass galaxies typically overestimated. Overall, this work suggests that we should be optimistic about our ability to infer the masses of high-z galaxies observed with JWST (notwithstanding contamination from AGN) but careful when modelling the impact of strong emission lines on broadband photometry.
Reference graph
Works this paper leans on
-
[1]
Adams, N. J., Bowler, R. A., Jarvis, M. J., Häußler, B., & Lagos, C. D. 2021, MNRAS, 506, 4933, doi: 10.1093/mnras/stab1956 Asplund, M., Grevesse, N., Sauval, A. J., & Scott, P. 2009, ARA&A,47,481,doi:10.1146/annurev.astro.46.060407.145222 Baldry, I. K., Driver, S. P., Loveday, J., et al. 2012, MNRAS, 421, 621, doi: 10.1111/j.1365-2966.2012.20340.x Barro,...
arXiv 2021
-
[8]
https://arxiv.org/abs/2001.06025 Katz, H., Rosdahl, J., Kimm, T., et al. 2023, OJA, 6, doi: 10.21105/astro.2309.03269 Kikuchihara, S., Ouchi, M., Ono, Y., et al. 2020, ApJ, 893, 60, doi: 10.3847/1538-4357/ab7dbe Kimm, T., & Cen, R. 2014, ApJ, 788, doi: 10.1088/0004-637X/788/2/121 Kroupa, P. 2002, Science, 295, 82 Labbé, I., van Dokkum, P., Nelson, E., et ...
work page Pith review arXiv 2001
-
[11]
https://arxiv.org/abs/2310.12228 Song, M., Finkelstein, S. L., Ashby, M. L. N., et al. 2016, ApJ, 825, 5, doi: 10.3847/0004-637x/825/1/5 Sorba, R., & Sawicki, M. 2015, MNRAS, 452, 235, doi: 10.1093/mnras/stv1235 —. 2018, MNRAS, 476, 1532, doi: 10.1093/mnras/sty186 Stanway, E. R., & Eldridge, J. J. 2018, MNRAS, 479, 75, doi: 10.1093/mnras/sty1353 Steinhard...
work page Pith review arXiv 2016
-
[13]
https://arxiv.org/abs/arXiv:2308.09665v1 Walcher, J., Groves, B., Budavári, T., & Dale, D. 2011, Fitting the integrated spectral energy distributions of galaxies, doi: 10.1007/s10509-010-0458-z Weaver, J. R., Davidzon, I., Toft, S., et al. 2023, A&A, 677, A184, doi: 10.1051/0004-6361/202245581 Weibel, A., Oesch, P. A., Barrufet, L., et al. 2024, eprint ar...
work page Pith review arXiv 2011
-
[37]
https://arxiv.org/abs/2310.08829 Cochrane, R. K., Best, P. N., Smail, I., et al. 2021, MNRAS, 503, 2622 Cochrane, R. K., Hayward, C. C., Anglés-Alcázar, D., et al. 2019, MNRAS, 488, 1779, doi: 10.1093/mnras/stz1736 Cochrane, R. K., Anglés-Alcázar, D., Mercedes-Feliz, J., et al. 2023, MNRAS, 523, 2409, doi: 10.1093/mnras/stad1528 Cole,S.,Norberg,P.,Baugh,C...
work page Pith review arXiv 2021
-
[128]
N., Kondapally, R., Williams, W
https://arxiv.org/abs/2305.14418 Best, P. N., Kondapally, R., Williams, W. L., et al. 2023, MNRAS, 523, 1729, doi: 10.1093/mnras/stad1308 Bhatawdekar, R., Conselice, C. J., Margalef-Bentabol, B., & Duncan, K. 2019, MNRAS, 486, 3805, doi: 10.1093/mnras/stz866 Bisigello, L., Caputi, K. I., Colina, L., et al. 2019, ApJS, 243, 27, doi: 10.3847/1538-4365/ab291...
arXiv 2023
-
[141]
https://arxiv.org/abs/2212.01915 Panter, B., Jimenez, R., Heavens, A. F., & Charlot, S. 2007, MNRAS, 378, 1550, doi: 10.1111/j.1365-2966.2007.11909.x Papovich, C., Dickinson, M., & Ferguson, H. C. 2001, ApJ, 559, 620, doi: 10.1086/322412 Papovich, C., Cole, J. W., Yang, G., et al. 2023, ApJL, 949, L18, doi: 10.3847/2041-8213/acc948 Pforr, J., Maraston, C....
arXiv 2007
-
[417]
2024, MNRAS, 529, 3751, doi: 10.1093/mnras/stae776 Ciesla, L., Elbaz, D., Ilbert, O., et al
https://arxiv.org/abs/1903.11082 Choustikov, N., Katz, H., Saxena, A., et al. 2024, MNRAS, 529, 3751, doi: 10.1093/mnras/stae776 Ciesla, L., Elbaz, D., Ilbert, O., et al. 2024, A&A, 686, A128, doi: 10.1051/0004-6361/202348091 Cochrane, R. K., Anglés-Alcázar, D., Cullen, F., & Hayward, C. C. 2024, ApJ, 961,
arXiv 1903
Show all 13 references
-
[1000]
https://arxiv.org/abs/0309134v1 Buat, V., Heinis, S., Boquien, M., et al. 2014, A&A, 561, 1, doi: 10.1051/0004-6361/201322081 Byler,N.,Dalcanton,J.J.,Conroy,C.,&Johnson,B.D.2017,ApJ, 840, 44, doi: 10.3847/1538-4357/aa6c66 Calzetti, D., Armus, L., Bohlin, R., et al. 2000, ApJ, ...
2014
-
[2020]
G., Förster Schreiber, N
https://arxiv.org/abs/2006.03599 Marchesini, D., Van Dokkum, P. G., Förster Schreiber, N. M., et al. 2009, ApJ, 701, 1765, doi: 10.1088/0004-637X/701/2/1765 Matthee, J., Naidu, R. P., Brammer, G., et al. 2023, eprint arXiv:2306.05448. https://arxiv.org/abs/2306.05448 McLeod, D...
2006 arXiv
-
[2023]
2023, Nature Astronomy, 7, 731, doi: 10.1038/s41550-023-01937-7 Bruzual, G., & Charlot, S
https://arxiv.org/abs/2303.00306 Boylan-Kolchin, M. 2023, Nature Astronomy, 7, 731, doi: 10.1038/s41550-023-01937-7 Bruzual, G., & Charlot, S. 2003, MNRAS, 344,
2023 arXiv
-
[3563]
https://arxiv.org/abs/2305.04944 Trussler, J. A. A., Conselice, C. J., Adams, N., et al. 2023, 23,
2023 arXiv
-
[3828]
J., Mortlock, A., et al
https://arxiv.org/abs/arXiv:1910.07524v1 High-z stellar masses can be recovered robustly with JWST photometry 15 Duncan, K., Conselice, C. J., Mortlock, A., et al. 2014, MNRAS, 444, 2960, doi: 10.1093/mnras/stu1622 Ferland, G. J., Chatzikos, M., Guzmán, F., et al. 2017, Revist...
1910 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.