{"id":"ca17a6b0-a4cc-447c-a5ce-73b8da385281","arxiv_id":"2509.01670","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Observed 1.5-2.5 micron excess in embedded young star clusters in the FEAST galaxies that CIGALE stellar population models fail to fit, strongest for the youngest and lowest-mass clusters.","lead":"JWST and Hubble images of young, still-cloud-wrapped star clusters in four nearby galaxies show a persistent excess of light at 1.5-2.5 microns that current stellar population models cannot reproduce. That mismatch suggests ages and masses measured for these embedded clusters with standard SED fitting are partly wrong, and that models need new ingredients such as dusty disks around newborn stars and statistical sampling of small clusters.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Nebular-grid coverage, not stellar ingredients, could explain the reported 1.5–2.5 μm excess; the tested insensitivity is only internal to the grid.","rationale":"The reader's weakest_assumption identifies exactly the condition that would invalidate the central claim: if the nebular grid does not bracket the true ionized-gas emission, the reported 'unaccounted NIR excess' is a template-coverage effect rather than evidence for missing stellar/population ingredients. The paper's own tests show insensitivity only within the tested grid, not completeness of the grid. This is more load-bearing than the age/mass circularity, which affects the secondary characterization of the excess rather than its existence, and more directly relevant than the dust-grid concern, which the authors already address with dust-free fits in Sec. 5.2. A Paα-based prediction of the nebular continuum is a decisive, quantitative check: the Paα line flux is measured in the same data, so the expected free-free/free-bound contribution at 1.5–2.0 μm can be computed without invoking new templates. If the residual exceeds that prediction, the central claim survives; if not, the conclusion should be reframed as a nebular-model completeness issue. The current analysis is careful and the excess is plausibly real, so no rejection is warranted; the reader's CONDITIONAL verdict remains appropriate.","tokens_in":30025,"tokens_out":7904,"duration_ms":104757,"concrete_test":"Use the measured Paα line fluxes (from the F187N continuum-subtracted maps) for the low-mass, young subsample to compute the expected nebular free-free + free-bound continuum at 1.5 and 2.0 μm with standard recombination theory (e.g., Osterbrock & Ferland), including the adopted extinction and the full allowed f_esc/f_dust range. Compare this predicted continuum flux to the observed residual flux in F150W and F200W. If the residual is consistent with the maximum predicted nebular continuum (especially for f_esc < 0.01), the excess is a nebular-grid coverage artifact; if it exceeds the maximum prediction by more than a factor of two across the sample, the missing-stellar/YSO interpretation is supported. A complementary check is to re-fit a subset of eYSCIs with CIGALE using f_esc extended to 0.001 and log U extended to −1.5, and see whether the F150W/F200W residuals persist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the F150W/F200W excess is \"not accounted for by current stellar population models\"—rests on the assumption that the CIGALE model grid brackets the true eYSC SED. The weakest point is the nebular grid (Sec. 3.1.1, Tab. 3). The tested ranges are log U = −3.5…−2, n_e = 10, 100 cm^-3, f_esc = 0.01…0.6, and f_dust = 0.01…0.3. The paper's robustness test only varies parameters within these ranges and finds insensitivity; it does not demonstrate that the ranges include the physically correct values for embedded, compact, very young clusters. At 1.5–2.5 μm, free-free and free-bound nebular continuum can be strong in dense HII regions. If the true nebular continuum is stronger than any grid model (e.g., f_esc effectively below 0.01, higher n_e, or different geometry/covering factor), the F150W/F200W residuals would be explained by nebular emission rather than by missing stellar populations or YSOs. Because the headline conclusion interprets the excess as evidence for missing stellar/population ingredients, this grid-coverage assumption is load-bearing and is not conclusively tested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper combines HST and JWST/NIRCam photometry to fit 0.2–5 μm SEDs of ~3800 emerging young star clusters (eYSCs) in four FEAST galaxies (M51, M83, NGC 628, NGC 4449) with CIGALE. The central empirical claim is a systematic model-observation residual at 1.5–2.5 μm (F150W and F200W), where CIGALE best-fit fluxes underestimate the observed fluxes; the residual is reported for all four galaxies and appears strongest for clusters with CIGALE-fitted ages ≤6 Myr and masses ≤3000 M⊙. The authors test and reject aperture size and dust-model choice as the origin, show that a dust-free fit up to F200W preserves the excess, and use slug stochastic IMF simulations to show that PMS stars plus extinction can partially but not fully reproduce the scatter. They conclude that current SSP-based models are missing one or more NIR ingredients, possibly YSO emission, stochastic IMF/PMS contributions, or non-homogeneous dust geometries. The paper also reports that CIGALE assigns ages ≥6 Myr to a fraction of strong Paα emitters whose Paα equivalent widths suggest younger ages.","tokens_in":30227,"tokens_out":5820,"duration_ms":71671,"significance":"If the central claim holds, this is an important empirical falsification test for current SED models of embedded young clusters: a reproducible 1.5–2.5 μm excess across four galaxies and hundreds to thousands of clusters would affect age/mass derivations and motivate inclusion of YSO SEDs, PMS emission, and stochastic IMF sampling in cluster SED fitting. The paper has clear strengths: the residual is quantified consistently across multiple environments; the aperture test (App. B), the dust-free test (Sec. 5.2), and the slug comparison (Sec. 5.4) are good robustness checks; and the public data products are a useful community resource. However, the headline claim that the excess is 'not accounted for by current models' is stronger than what the tested CIGALE grid demonstrates, because only a restricted nebular parameter cube is explored. In addition, the demographic statement that the excess is strongest in young, low-mass clusters is made using age and mass bins from the very fits that fail in the NIR, which introduces a partial circularity. These issues are addressable, but they are load-bearing for the paper's main conclusions.","major_comments":[{"comment":"The claim that the F150W/F200W excess is 'not accounted for by current stellar population models' is too strong relative to the experiments. The tested nebular grid is log U = -3.5…-2, n_e = 10, 100 cm^-3, f_esc = 0.01…0.6, f_dust = 0.01…0.3. The insensitivity test in Sec. 3.1.1 varies parameters only within these ranges; it does not establish that the ranges bracket the true nebular emission of compact, embedded, very young HII regions. At 1.5–2.5 μm, free-free and free-bound nebular continuum can be significant, and a stronger continuum than any grid point (e.g., higher n_e, very low f_esc, or different covering factor) would produce exactly this residual pattern. Please either extend the grid to physically motivated extremes (e.g., n_e up to 10^4 cm^-3, f_esc below 0.01, larger f_dust), quantify the maximal nebular contribution at F150W/F200W, or temper the Abstract/Sec. 6 wording to","section":"Sec. 3.1.1 / Tab. 3"},{"comment":"The statement that the NIR excess is 'most prominent in low-mass (≤3000 M⊙) and young (≤6 Myr) clusters' is based on bins of best-fit age and stellar mass from the same CIGALE fits that fail in the NIR. A missing NIR component can bias both fitted ages and masses, so the demographic conclusion is partially circular. What is robust is the direct model-observation residual itself; the dependence of that residual on fitted parameters is a secondary, model-dependent result. The Paα EW analysis in Sec. 5.6 addresses only the misclassification of the >6 Myr bin and is itself affected by excess continuum in F150W/F200W. To support the headline trend, use an independent age or mass indicator (e.g., Paα EW, HST/optical colors, or slug prior predictions), or explicitly rescale the claim to 'clusters that CIGALE classifies as young and low mass'.","section":"Sec. 4 / Fig. 6"},{"comment":"The choice of the 'adopted' dust grid is justified by lower reduced χ² rather than by physical priors, and the paper interprets the preference for low U_min / high γ as evidence of model shortcomings at 3–5 μm. Because the dust parameters are essentially unconstrained without MIR/FIR data, this preference could be a consequence of the same missing NIR component being absorbed by flexible dust templates rather than a separate physical failure. The dust-free fit in Sec. 5.2 is a good test for the 1.5–2.5 μm excess and supports that result, but it does not establish the separate claim about the dust parameter grid. I recommend presenting the dust-grid preference as a suggestive byproduct of the NIR excess, not as an independent demonstration of missing model components.","section":"Sec. 3.1.1 / Sec. 5.2"}],"minor_comments":[{"comment":"Typographical issues to correct: 'FULLBOX 4TIGHT' appears to be a malformed dither pattern name, and 'V arun' in the author list should be 'Varun'. These do not affect the science.","section":"Sec. 2.2"},{"comment":"The Paα EW formula mixes line flux and bandpass notation. Please clarify whether Fλ,line and Fλ,cont are in the same units and state how the F187N filter width (240 Å) is applied. Propagating photometric uncertainties into EW uncertainties would make the age comparison in Fig. 10 more informative.","section":"Sec. 5.6 / Eq. (1)"},{"comment":"The test of nebular-grid insensitivity in Appendix A is performed on optically selected YSCs in NGC 4449 using Hα, not on eYSCs in the NIR bands where the excess is found. This is a useful consistency check, but it is not equivalent to testing grid coverage at 1.5–2.5 μm for embedded clusters. The text should state this limitation explicitly.","section":"Appendix A"},{"comment":"The F814W residual distribution shows a negative tail that is described only as 'the opposite behavior'. A brief quantitative statement (e.g., median Δm and fraction of objects beyond ±0.5 mag) would help the reader assess whether this is a separate model deficiency or a compensating effect of the NIR excess.","section":"Fig. 5 / Sec. 4"},{"comment":"The yggdrasil evolutionary track is shown as a gray line, but the caption does not state all model assumptions (e.g., f_cov = 0.5, E(B-V)=0). Adding these values to the caption would improve reproducibility.","section":"Fig. 9 / Fig. 15"}],"recommendation":"major_revision","confidential_remarks":"The central residual is likely real and the paper is a useful contribution, but the 'missing ingredients' interpretation is not fully secured because the nebular grid coverage is the key untested assumption. The authors can address this with targeted extreme-grid tests and by softening the wording. I do not see grounds for rejection, but the claim as currently phrased is too strong."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Alex,\n\nThe paper does something useful: it quantifies, across ~3800 eYSCs in four galaxies, a systematic excess at 1.5–2.5 μm that standard SSP models fitted through CIGALE fail to reproduce. The residual is consistent, grows in the younger/lower-mass bins, and survives checks against aperture size, dust grid choice, and dust-free fits. That is a genuine, reproducible result and likely the first multi-galaxy documentation of this effect. The slug-based exploration of stochastic IMF sampling plus PMS stars is a thoughtful addition.\n\nThe main soft spot is the one the stress-test names: the nebular grid. The authors say results are insensitive to log U, n_e, f_esc, and f_dust within their tested ranges, but they never demonstrate that those ranges bracket reality for dense, embedded, very young clusters. If the true nebular continuum is stronger than any template in the grid, some or all of the F150W/F200W residual could be nebular rather than stellar or YSO emission. That is a load-bearing caveat for the headline \"missing ingredients\" claim. I don't think it sinks the paper—the excess is large, and the mass/age trends are not obviously nebular—but it should be stated and ideally tested with a wider grid or independent constraints.\n\nThe other caveat is more minor: the age/mass decomposition uses the same CIGALE fits that are failing, so saying the excess is strongest in low-mass young clusters is partly circular. The authors acknowledge this, and the Pa-alpha EW analysis partially corroborates, but it limits how strongly they can claim the mass dependence.\n\nThe dust grid choice is disclosed as chi2-driven rather than physically motivated, which is honest and sensible given the data. The paper would be stronger if it pushed the nebular grid further and tried a few YSO templates, but it is already a careful, data-rich contribution.\n\nWho should read it: anyone doing SED fitting of embedded clusters with JWST in FEAST, PHANGS, or similar surveys. It is a useful warning that current models miss something in the NIR, and the catalog release will be valuable. I'd send it to a good referee; the measurement likely stands even if the interpretation needs trimming.","headline":"Solid multi-galaxy measurement of a real NIR excess in young embedded clusters; the 'missing ingredients' interpretation is plausible but the nebular grid coverage caveat is not fully closed.","tokens_in":30988,"tokens_out":2667,"would_cite":true,"duration_ms":29113,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic flux excess at 1.5–2.5 microns, strongest in young (≤6 Myr) low-mass (≤3000 solar masses) clusters, is not captured by current stellar population models; stochastic IMF sampling plus pre-main-sequence stars shrinks but does not","keywords":["Star clusters","Spectral energy distribution","Near-infrared excess","Embedded young star clusters","Pre-main-sequence stars","Stochastic IMF sampling","Polycyclic aromatic hydrocarbons","JWST NIRCam"],"falsifier":"Take spectra of about ten low-mass (≤3,000 solar masses) clusters with SED-fitted ages ≤3 Myr in M51 over 1.4–2.6 µm with JWST/NIRSpec. If the observed continuum matches a stellar-plus-nebular model within a few percent band by band, the excess is a photometric or modeling artifact; if a broad smooth excess remains above the model and is not Pa-alpha, the missing component is physical.","tokens_in":29819,"feed_emoji":"🔭","tokens_out":9543,"duration_ms":104903,"temperature":0.7,"pith_summary":"This paper uses combined HST and JWST photometry from 0.2 to 5 microns to fit the spectral energy distributions of about 3,800 emerging young star clusters in four nearby galaxies and claims that current stellar population models systematically underestimate their 1.5–2.5 micron flux. The excess is largest in the youngest (≤6 Myr) and least massive (≤3,000 solar masses) clusters, appears in the F150W and F200W (and F277W) bands, and survives tests with different apertures, different dust grids, and fits with dust emission removed. Stochastic sampling of the initial mass function with pre-main-sequence stars in the slug code moves model colors toward the observed ones but cannot close the gap in the metal-rich spirals. If the claim holds, SED-fit ages and masses of embedded young clusters are biased, and a fraction of strong Pa-alpha emitters are misclassified as older than 6 Myr. The paper concludes that emission from young stellar objects and more realistic IMF sampling are the missing ingredients.","feed_headline":"Young cluster models miss a near-infrared excess at 1.5–2.5 µm","feed_subtitle":"Peaking in clusters below 3,000 solar masses and under 6 Myr, the excess biases SED-fit ages of embedded clusters.","key_machinery":"The load-bearing quantity is the per-filter residual Δm = m_model − m_observed, computed for thousands of eYSCs and summarized as medians in age/mass bins; positive Δm in F150W/F200W defines the NIR excess. The argument is carried by two SED tools: CIGALE, a deterministic fitting code using instantaneous-burst stellar populations, CLOUDY nebular emission, an attenuation law, and dust templates, and slug, a stochastic stellar-population synthesis code that samples the IMF and includes pre-main-sequence stars. The paired Pa-alpha equivalent-width age comparison provides an independent clock that exposes the age bias.","core_discovery":"The central discovery is a systematic near-infrared excess: observed fluxes in the 1.5 and 2.0 micron JWST bands exceed the best-fit model fluxes by roughly 0.2–0.6 mag for the young, low-mass subpopulation, while neighboring blue and red bands are fit well. The excess is strongest for clusters fitted with ages ≤3 Myr and masses ≤3,000 solar masses, and in every galaxy it persists when dust emission is removed from the fit and when smaller apertures are used. As a consequence, the fitting code assigns ages ≥6 Myr to a subset of clusters whose Pa-alpha equivalent widths indicate much younger ages, and it prefers dust parameters that are not those expected from mid/far-infrared studies of star","pith_inferences":["If optically selected young clusters already show ~0.2 mag residuals, legacy HST-based cluster catalogs may carry a milder version of the same NIR bias, systematically underestimating the NIR flux of their youngest objects.","The paper does not test whether the excess tracks accretion; correlating the 1.5–2.5 µm residual with Br-gamma or Pa-alpha line width within the ≤3 Myr bin would separate a YSO-disk origin from a purely photospheric PMS origin.","Adding YSO SED templates to the fitting grid is a direct falsification test: the excess should disappear and the fitted dust parameters should return to the Umin~1–10, gamma~0.5–1 ranges expected from MIR-FIR studies.","Because the 3.3 µm PAH continuum subtraction assumes stellar plus nebular plus dust model continua, a missing YSO component would bias measured PAH equivalent widths, with consequences for PAH-based star-formation-rate calibrations in embedded regions."],"forward_implications":["Ages from CIGALE for young, low-mass eYSCs are biased: a fraction of strong Pa-alpha emitters with equivalent widths pointing to ages well below 6 Myr are fitted with ages ≥6 Myr.","Any SED-fitting pipeline that uses deterministic, fully sampled IMFs with no pre-main-sequence stars will underestimate the 1.5–2.5 µm flux of embedded clusters and will push dust parameters to unrealistic values to compensate.","The Pa-alpha equivalent-width age relation is unreliable for clusters below roughly 5,000 solar masses, where stochastic sampling makes a given EW consistent with both young and old ages.","Adding young stellar object SEDs and stochastically sampled IMFs with pre-main-sequence emission to fitting codes is necessary to recover credible ages, masses, and dust parameters for emerging clusters."],"supporting_citations":[{"why":"Supplies the CIGALE SED-fitting code whose best-fit residuals define the near-infrared excess.","marker":"Boquien et al. 2019"},{"why":"Supplies the stellar population models used in the fits, the component that underestimates the 1.5–2.5 µm flux.","marker":"Bruzual & Charlot (2003)"},{"why":"Provides the CLOUDY nebular emission templates in the fitting grid.","marker":"Ferland et al. 2013"},{"why":"Provides the nebular continuum models that set the ionized-gas baseline in the grid.","marker":"Inoue 2011"},{"why":"Supplies the slug stochastic population synthesis code used to test IMF sampling and PMS-star emission.","marker":"da Silva et al. 2012"},{"why":"Provides the stellar tracks that include pre-main-sequence stars and rotation in the slug models.","marker":"Dotter 2016"},{"why":"Provides the deterministic yggdrasil SSP tracks used as comparison in the color-color diagrams.","marker":"Zackrisson et al. 2011"},{"why":"Supplies the starburst attenuation law applied to the stellar continuum in the fits.","marker":"Calzetti et al. 2000"},{"why":"Supplies the dust emission templates whose parameter ranges the fits stretch to compensate for the excess.","marker":"Draine et al. 2014"}],"fun_headline_variants":["Hidden NIR excess in young clusters skews age estimates","JWST shows NIR excess models can't fit in young clusters","1.5–2.5 µm excess in young, low-mass clusters puzzles models","Young cluster NIR glow exposes SED model gaps"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the grid of ionized-gas models used in the fitting covers all realistically strong nebular continuum at 1.5–2.5 microns; if the true nebular light is stronger than any grid entry, the reported excess is a limitation of the grid, not a missing stellar or dust component.","fun_headline_variants_meta":{"raw":{"variants":["Hidden NIR excess in young clusters skews age estimates","JWST shows NIR excess models can't fit in young clusters","1.5–2.5 µm excess in young, low-mass clusters puzzles models","Young cluster NIR glow exposes SED model gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000297,"raw_usage":{"total_tokens":1642,"prompt_tokens":913,"completion_tokens":729,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":657,"completion_tokens_details":{"reasoning_tokens":655}},"tokens_in":657,"tokens_out":729,"duration_ms":9179,"temperature":1.0,"reasoning_tokens":655,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:18:05.030557+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take spectra of about ten low-mass (≤3,000 solar masses) clusters with SED-fitted ages ≤3 Myr in M51 over 1.4–2.6 µm with JWST/NIRSpec. If the observed continuum matches a stellar-plus-nebular model within a few percent band by band, the excess is a photometric or modeling artifact; if a broad smooth excess remains above the model and is not Pa-alpha, the missing component is physical.","supporting_citations":[],"review_version":1}