{"id":"0f932bef-941b-453f-885c-2f1aa3031858","arxiv_id":"2608.09439","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Young star cluster mass functions have a nearly universal M^-2 slope, but their high-mass truncations vary by orders of magnitude across and within galaxies, with no simple correlation to star formation rate or shear.","lead":"Astronomers fit the masses, ages, and extinction of 8,276 star clusters in 12 nearby galaxies, using a neural network to correct for which clusters a survey would miss. The result challenges simple recipes that link the largest cluster masses to a galaxy's star formation rate or rotation shear.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The sub-galactic M_break variations could be completeness artifacts: c-4 is trained at field level and applied to radial subregions without validation that the selection function is position-independent.","rationale":"The paper is a careful, well-documented forward-modelling study with a large homogeneous sample, public code, and posterior predictive checks that reproduce the UV and blue luminosity functions well. The central galaxy-wide result, that M_break is often broad, multimodal, and not simply tied to galaxy-averaged Sigma_SFR, is supported by the full posterior treatment and by the explicit discussion of the alpha_M-M_break covariance. The most fragile link in the chain is the transfer of the field-level c-4 completeness model to radial subregions. The sub-galactic M_break contrasts in NGC 5457 and NGC 5194-5195 are the strongest evidence for environmental dependence, but they are exactly the measurements most sensitive to a position-dependent selection function. Because c-4 conditions on photometry alone, and the training injections are distributed over the whole field, the model averages over crowding and background conditions; radial bins sample different parts of that distribution. Without a validation that the recovery fraction at fixed magnitude is stable across the radial bins, the reported inner/outer differences could be produced by completeness gradients rather than by genuine changes in the cluster mass function. The reader's weakest_assumption identified precisely this concern, and the F814W systematic is a secondary worry that does not specifically threaten the sub-galactic comparison. A targeted recovery-fraction split and refit would settle the issue. If the differences survive, the conclusion is strengthened; if they do not, the central environmental claim would need to be restricted to galaxy-wide scales. This does not move the verdict, because the reader already marked the paper CONDITIONAL for exactly this reason.","tokens_in":29328,"tokens_out":2808,"duration_ms":30817,"concrete_test":"Evaluate or retrain c-4 separately in each radial bin used in Table 3: take the existing artificial-cluster injection outcomes, split them by the same R_gal divisions, and compare the recovery fraction versus magnitude in each bin with the field-level c-4 prediction used in the fits. Then refit the MID model for NGC 5457, NGC 5194-5195, and NGC 628 using bin-specific bP_obs. If the inner/outer M_break differences in Table 3 persist within the posterior widths, the environmental signal survives; if they shrink or vanish, the sub-galactic claim is a completeness artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The weakest load-bearing assumption is in Section 3.2: the c-4 completeness model is trained per LEGUS field from artificial cluster tests whose injection positions follow the observed galaxy light distribution, but it is then applied to radial subdivisions and individual fields without retraining or testing. Since c-4 returns bP_obs,F(m_j) as a function of photometry only, it implicitly marginalizes over the spatial distribution of crowding and background within the whole field. Radial bins are exactly where crowding, surface brightness, and hence recovery probability vary most. If the average completeness within a radial bin differs from the field-average at fixed magnitude, the reweighted likelihood in Equation 4 will misestimate the expected number of faint/low-mass clusters relative to bright/high-mass ones, directly shifting the inferred Schechter slope and M_break. Table 3's dramatic inner/outer M_break differences, for example NGC 5457 with log10(M_break/M_sun) = 7.08 inner versus 4.37 outer, are the central evidence for environmental dependence, so this is not a peripheral technicality. The paper states that 'we expect the galactic-level completeness model to provide an adequate approximation for the radial subsamples', but supplies no recovery-fraction comparison for radial bins. The F814W residual discussed in Section 5.3 is a second systematic, but it affects galaxy-wide and sub-galactic fits similarly; the radial-transfer assumption is what specifically underpins the sub-galactic environmental claim. This concern does not by itself overturn the qualitative conclusion, but it is the point that must be settled before the sub-galactic result can be accepted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a Bayesian forward-modelling analysis of ~8300 star clusters in 12 LEGUS galaxies, using the slug stellar population synthesis framework and a neural-network completeness estimator (c-4) to infer cluster mass functions, age distributions, and extinction distributions at both galaxy-wide and sub-galactic scales. The main claims are that cluster mass function slopes are broadly consistent with a power law of slope near -2, that the Schechter truncation mass M_break varies by orders of magnitude between and within galaxies, that mass-independent disruption is favoured over mass-dependent disruption, and that these variations are not well described by simple one-parameter environmental scalings such as star formation rate surface density or shear. The paper also reports posterior predictive checks that reproduce observed luminosity functions in four of five broad bands, with a residual in F814W that the authors attribute to stellar population synthesis modelling.","tokens_in":29621,"tokens_out":6067,"duration_ms":65845,"significance":"If the main conclusions hold, this is an important step in star cluster demographics: it demonstrates that a full forward-modelling treatment with a continuous selection function can extract demographic parameters from partially incomplete catalogues, and it challenges the existence of a universal relation between galaxy-averaged star formation rate surface density and the cluster mass function truncation mass. The paper is explicit about its methodological debt to Tang et al. (2026) for the c-4 completeness model, and it provides posterior predictive checks and openly acknowledges the F814W systematic. The sub-galactic comparison with the Reina-Campos & Kruijssen (2017) framework is a useful, falsifiable confrontation. The main weakness is that the radial subregion analysis rests on an unvalidated transfer of field-level completeness to radial bins, which is directly relevant to the paper's strongest claim of sub-galactic environmental dependence.","major_comments":[{"comment":"The sub-galactic analysis assumes that field-level c-4 completeness models can be applied to radial subregions without retraining, but this assumption is load-bearing and untested. The quantity bP_obs,F(m_j) is learned from artificial cluster tests whose injection positions are drawn from the observed galaxy light distribution, so it marginalizes over the spatial distribution of crowding and background within the whole field. A radial bin samples only a subset of those conditions, and the average recovery probability at fixed magnitude in the bin need not equal the field average. Because Eq. (4) multiplies the intrinsic population weight by bP_obs,F(m_j), a radial-dependent offset in completeness directly biases the relative weights of faint compared to bright clusters and hence the inferred alpha_M and M_break. The paper states in Section 3.2 that 'we expect the galactic-level completeness model to provide an adequate approximation for the radial subsamples', but it supplies no empirical check of this expectation. This matters because the inner/outer differences in Table 3, for example log10(M_break/M_sun)=7.08 versus 4.37 for NGC 5457, are the central evidence for sub-galactic environmental dependence. Please validate the assumption by computing recovery fractions as a function of magnitude in the radial bins from the existing artificial-cluster tests, or retrain c-4 for the subregions and compare the inferred parameters.","section":"Section 3.2, Eq. (4)"},{"comment":"The paper acknowledges a systematic residual in the F814W band for all three representative galaxies, but it does not propagate this systematic into the reported demographic parameter constraints. The text states that the residual 'can shift the inferred masses, ages, and extinctions of individual clusters' and that the implication is 'non-negligible', yet the posterior intervals in Tables 2 and 3 are computed from a likelihood that does not include a model for this residual. Since the central claims about M_break variation and the comparison with the Johnson et al. (2017) and Reina-Campos & Kruijssen (2017) relations depend on the fitted M_break values, the quoted intervals are formally too narrow. Please quantify the impact of the F814W residual, for example by refitting the sample without F814W or by adding a band-dependent nuisance term, and state explicitly whether the main conclusions (the orders-of-magnitude variation in M_break and the disagreement with the Sigma_SFR-M_break relation) survive.","section":"Section 5.3, Fig. 10"},{"comment":"The conclusion that the Reina-Campos & Kruijssen (2017) model 'fails to reproduce' the observations is weakened by the way the model uncertainty envelope is constructed. In Appendix A, the upper bound of the model envelope is defined as the maximum of the present-day and historical predictions, where the historical prediction uses an ad hoc factor f_hist that is chosen per region (e.g., f_hist=10 for NW2 and NW3). This makes the model comparison effectively partially fitted to the data, since the choice of f_hist is motivated by the same observational evidence that the comparison is intended to test. The paper should either present the comparison with a prespecified f_hist, show the sensitivity of the conclusion to f_hist, or soften the claim that the model cannot explain the observed radial trends.","section":"Section 5.2.4 and Appendix A, Eqs. (A12)-(A13)"}],"minor_comments":[{"comment":"The text says 'span nearly one orders of magnitude'; this should be 'span nearly an order of magnitude'.","section":"Section 2.1"},{"comment":"The convergence criterion 'N_post >= 50 tau_int' is described with 'our 1500 post-burn-in iterations', but with 100 walkers the effective number of post-burn-in samples is 150,000; please clarify whether the criterion is applied per walker or to the pooled sample.","section":"Section 3.4"},{"comment":"The sub-galactic analysis reports only MID results and states that tests against MDD are omitted because MID is preferred, but no Akaike weights are given for the radial subregions; please quantify the MID preference for the subregions or state that model choice does not affect the CMF conclusions.","section":"Section 4.2"},{"comment":"The top-left panel of Figure 10 combines F435W and F438W in one column; the caption should state which galaxies are observed in each filter and whether the combined panel is a weighted average or a plot of one filter.","section":"Figure 10"},{"comment":"The table lists only marginal medians for alpha_M and M_break, but the text emphasizes that these parameters are strongly covariant; a joint posterior plot for the radial subdivisions would help the reader assess the significance of the reported differences.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of MNRAS and the methodology is largely sound, with honest treatment of the F814W residual. The central issue is the unvalidated transfer of field-level completeness to radial bins, which is directly load-bearing for the sub-galactic claims. If the authors can demonstrate robustness of the radial results to the completeness assumption, or refit with per-region completeness, the paper would be publishable. The F814W residual should be addressed quantitatively rather than only qualitatively. I do not see grounds for rejection, but the current form requires substantive revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper is worth taking seriously. It is the first multi-galaxy application of the c-4 neural completeness model inside the slug Bayesian forward-modelling framework, and the machinery is used honestly. The galaxy-scale results—CMF slopes near -2, no clean Sigma_SFR–M_break relation, MID preferred over MDD in 9/12 galaxies—are credible, and the paper's own caveats are clearly stated. The posterior predictive checks match observed luminosity functions in F275W–F555W for representative galaxies; the F814W residual is flagged and discussed as a stellar-population systematic. That is good practice.\n\nThe soft spot that matters is the radial completeness transfer. In Section 3.2 the authors say they do not retrain c-4 for radial subdivisions and expect the galactic-level completeness model to be adequate, but they do not test that expectation. Recovery probability is a strong function of crowding and background, both of which vary systematically with radius. If the field-averaged completeness is not representative of an inner or outer bin at fixed magnitude, Equation 4 misweights faint clusters relative to bright ones and can shift the inferred Schechter slope and M_break. Table 3's dramatic values (NGC 5457 log M_break 7.08 inner vs 4.37 outer) are the central evidence for sub-galactic environmental dependence, so this is not a peripheral technicality. I would not call it a fatal flaw: the galaxy-scale results and the central-field/outer-field differences may survive, but the quantitative sub-galactic M_break values should be treated as provisional until the authors produce a recovery-fraction comparison for the radial subsamples.\n\nMinor point: the sub-galactic fits lack the posterior-predictive validation that is shown for the galaxy-wide fits, so we only have the overall model check, not a check in the exact bins where the strongest claims live.\n\nOverall: a solid paper with one load-bearing assumption that is currently unverified. It deserves a serious referee; I would send it out, and the referee should ask for a radial completeness validation before acceptance. If I were working on cluster demographics, I'd cite the galaxy-scale results and the methodology, while waiting on the sub-galactic numbers.","headline":"A careful, useful forward-modelling study that makes a good case for galaxy-scale environmental complexity, but the sub-galactic M_break story needs a completeness check before I'd quote it.","tokens_in":30203,"tokens_out":4536,"would_cite":true,"duration_ms":48996,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Star-cluster demographics are environmentally dependent, but no one-parameter model currently explains the variation in the high-mass cutoff.","keywords":["star cluster demographics","cluster mass function","LEGUS survey","completeness","Bayesian forward modelling","cluster disruption","galactic environment","Schechter function"],"falsifier":"Retrain the c-4 completeness model separately on artificial clusters injected only in the inner ($R<2.4$ kpc) and outer regions of NGC 5457 and re-fit the mass function: if the inferred break masses converge, the reported environmental gradient is a catalogue-completeness artifact; if they stay apart, the environmental signal is real.","tokens_in":29092,"feed_emoji":"🌌","tokens_out":10661,"duration_ms":101910,"temperature":0.7,"pith_summary":"This paper tries to establish that the demographics of young star clusters depend on galactic environment in a way that is real but not simple. Using about 8300 clusters from 12 LEGUS galaxies and a forward-modelling approach that accounts for catalogue incompleteness, the authors find that the low-mass slope of the cluster mass function is near universal, close to a power law $M^{-2}$, while the high-mass truncation mass varies by orders of magnitude between and within galaxies. The variation does not track star-formation-rate surface density or shear, and the two theoretical models they test do not reproduce it. If right, this means the upper end of the cluster mass function is environmentally regulated by a combination of local conditions, not by any single galaxy-wide parameter.","feed_headline":"High-mass cluster cutoffs shift with environment","feed_subtitle":"Across 12 nearby galaxies, the low-mass cluster slope stays near -2 while the truncation mass wanders.","key_machinery":"The load-bearing tool is the c-4 neural-network completeness estimator, a model trained on artificial cluster injection tests to output the probability that a cluster with given photometric properties would be recovered by the survey pipeline. This probability feeds directly into the likelihood of a stochastic forward-modelling framework: synthetic clusters are drawn from a library, weighted by the assumed mass-age-extinction distribution and by their catalogue-inclusion probability, and the predicted photometric catalogue is compared with the observed one. This replaces hard completeness cuts and lets partially detected faint and low-mass clusters contribute to the inference. The comparison is done separately for each filter combination, and posterior sampling of the demographic parameters uses an ensemble MCMC sampler.","core_discovery":"The central discovery is a pattern in cluster demographics: the slope of the young cluster mass function is broadly universal, consistent with $M^{-2}$, but the mass at which the function truncates shifts by more than two decades between galaxies whose star-formation surface densities are within a factor of two of one another. Within single galaxies, sub-galactic regions also show differences: the central fields of NGC 5457 and NGC 5194-5195 have higher or effectively untruncated break masses while outer fields are truncated near $10^{4.3}$ to $10^{4.7}$ solar masses, whereas NGC 628 shows no radial break variation but differs between its observational fields. The paper also finds that the cluster age distribution is best described by mass-independent disruption in most of the 12 galaxies, with wide variation in when disruption begins. The authors conclude that environmental regulation of cluster formation is real but multivariate: none of the tested current models, whether shear-based, gas-and-dynamics-based, or a simple star-formation-surface-density relation, explains the observed pattern.","pith_inferences":["Editorial inference: the paper's own completeness caveat suggests a decisive control experiment - retrain the completeness model on radial subregions of NGC 5457; if the inner-outer break-mass difference persists, it is environmental, and if it shrinks, it is a detection artifact.","Editorial inference: the red-band residual reported in the paper will shift inferred masses and ages through the age-extinction degeneracy, so deeper and redder follow-up data are the most direct way to tighten constraints on the upper cluster-mass scale.","Editorial inference: the failure of one-parameter relations points toward stacking sub-galactic bins across many galaxies and jointly fitting local gas surface density, epicyclic frequency, pressure, and recent star-formation history as predictors of the break mass."],"forward_implications":["The low-mass slope of the cluster mass function is approximately universal, so any theory of cluster formation must explain why the slope stays near $-2$ across environments while the high-mass cutoff moves.","A galaxy-wide star-formation-rate surface density does not determine the break mass; NGC 628, for example, rules out truncation below about $10^6$ solar masses despite a star-formation surface density that would predict a much lower break.","Rotation-curve shear alone fails as a predictor: NGC 5457 and NGC 5194-5195 show inner-outer differences, but NGC 628 shows none across the same shear boundary.","Cluster disruption appears mass-independent over the observable range, with the onset time of disruption varying from about a million years to more than a gigayear among galaxies.","Reported break masses should be interpreted through full posterior distributions, because the cluster mass function slope and break mass are strongly covariant and single best-fit values can mislead."],"supporting_citations":[{"why":"Supplies the Bayesian forward-modelling framework and the mass-independent and mass-dependent disruption models used for inference.","marker":"Krumholz et al. (2019b)"},{"why":"Introduces the c-4 neural network completeness estimator that replaces hard completeness cuts.","marker":"Tang et al. (2026)"},{"why":"Defines the LEGUS survey and its cluster catalogue construction pipeline.","marker":"Calzetti et al. (2015)"},{"why":"Provides the LEGUS cluster catalogues and morphological classifications used as the observed data.","marker":"Adamo et al. (2017)"},{"why":"Proposes the empirical star-formation-surface-density versus break-mass relation that the paper tests and finds inconsistent.","marker":"Johnson et al. (2017)"},{"why":"Provides the rotation-curve shear framework and the beta equals 0.3 threshold used for sub-galactic divisions.","marker":"Suwannajak et al. (2014)"},{"why":"Gives the maximum-cluster-mass model whose predictions are compared with the fitted break masses.","marker":"Reina-Campos & Kruijssen (2017)"},{"why":"Provides the model profile and uncertainty band for NGC 5194-5195 used in the sub-galactic comparison.","marker":"Messa et al. (2018)"},{"why":"Supplies galaxy-wide star-formation surface densities and other environmental parameters for the sample.","marker":"Menon et al. (2021)"},{"why":"Compiles literature break-mass measurements against which the new posterior distributions are compared.","marker":"Wainer et al. (2022)"}],"fun_headline_variants":["Cluster mass cutoff wanders by decades across galaxies","Star cluster slope stays -2, cutoff varies wildly","No simple model explains cluster cutoff shifts","Environment shifts cluster truncation mass, not slope","Cluster demographics: universal slope, erratic cutoff"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The completeness model trained once per galaxy field is assumed to remain valid for smaller radial subregions of that field, even though crowding, background, and detection conditions vary within the field.","fun_headline_variants_meta":{"raw":{"variants":["Cluster mass cutoff wanders by decades across galaxies","Star cluster slope stays -2, cutoff varies wildly","No simple model explains cluster cutoff shifts","Environment shifts cluster truncation mass, not slope","Cluster demographics: universal slope, erratic cutoff"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000487,"raw_usage":{"total_tokens":2419,"prompt_tokens":981,"completion_tokens":1438,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":1369}},"tokens_in":597,"tokens_out":1438,"duration_ms":11683,"temperature":1.0,"reasoning_tokens":1369,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:20:40.774876+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the c-4 completeness model separately on artificial clusters injected only in the inner ($R<2.4$ kpc) and outer regions of NGC 5457 and re-fit the mass function: if the inferred break masses converge, the reported environmental gradient is a catalogue-completeness artifact; if they stay apart, the environmental signal is real.","supporting_citations":[{"cited_title":"Monthly Notices of the Royal Astronomical Society , volume =","cited_arxiv_id":null,"evidence_quote":"Introduces the c-4 neural network completeness estimator that replaces hard completeness cuts."}],"review_version":1}