{"id":"6cf5610c-285c-4f7d-9bbc-e36a846f700a","arxiv_id":"2608.07671","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In 42 dwarf galaxies, radial age gradients correlate strongly with global formation history in a way that favors simulations without radially breathing gas flows.","lead":"The paper measures how the average age of stars changes from the center to the edge of 42 nearby dwarf galaxies, using Hubble images to reconstruct each galaxy's star formation history in concentric rings. It finds the pattern of aging correlates with the galaxy's overall formation history, and this lets observations tell apart competing computer simulations of dwarf galaxy evolution.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The gamma-90 vs tau-90(global) correlation reuses the same four-point radial fit for both axes; the 86% validation covers only complete-coverage galaxies, leaving the 29% incomplete-azimuthal subsample unvalidated.","rationale":"The reader's weakest assumption identifies exactly the load-bearing step: tau_90(global) is not an independent galaxy-wide measurement but the same linear fit evaluated at R_hl. I read the full text and the correlation analysis carefully. The most statistically significant claim (gamma_90 vs tau_90(global), p<0.001) depends on this construction, and the validation in Section 3.3 is limited to complete-coverage targets and passes in only 86% of those cases. For the 29% of targets with incomplete azimuthal coverage, there is no direct validation that the interpolated value equals a true global tau. Because gamma_90 is the slope of the same fit, any systematic bias in the fit affects both axes coherently, making the reported p-value untrustworthy until this is tested. The gamma_50 null result is more robust to this concern because the induced covariance is positive, so it cannot explain a null; the simulation-discrimination claim therefore survives this particular concern. I also considered the non-volume-limited sample and the differences between simulation definitions of tau, but these are acknowledged and do not threaten the central correlation as directly. The proposed test on the A=1 subset would settle whether the headline correlation persists when the global quantity is measured independently. Since the paper already receives a CONDITIONAL verdict, this concern does not change the verdict.","tokens_in":52260,"tokens_out":4918,"duration_ms":55825,"concrete_test":"Restrict to the 30 A=1 targets and recompute tau_90(global) independently by fitting a single MATCH SFH to all stars within 4.4 h_r (the sum of the four complete annuli), rather than interpolating the radial fit at R_hl. Rerun the Pearson and Spearman correlations of gamma_90 against this directly measured tau_90(global). If the p-value rises above 0.05 or the correlation coefficient drops by more than ~0.2, the headline correlation is at least partly an artifact of reusing the same fitted line. As a secondary check, for the 12 A<1 targets, compute a coverage-corrected global tau by area-weighting the four annulus SFHs and recompute the full-sample correlation; if the coefficients shift materially, incomplete azimuthal coverage is a driving bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 defines gamma_90 as the slope of a linear fit to per-radial-bin tau_90 values, then defines tau_90(global) by evaluating that same best-fit line at R_hl. Both headline quantities are therefore linear combinations of the same four fitted tau_90 values, not independent measurements. Correlated bin-level errors or a biased fit (e.g., from incomplete azimuthal coverage) shift the slope and the R_hl intercept coherently, which can inflate the reported p<0.001 gamma_90-tau_90(global) correlation even if the true physical correlation is weaker. The paper's validation that interpolation at R_hl reproduces directly measured global values in 86% of cases is restricted to the A=1 subset (30/42 galaxies); for the 12 targets with A<1 (29% of the sample), no direct check is possible, and the same fitted line is the only source of the 'global' value. The gamma_50 null result is less vulnerable to this particular artifact, since the induced covariance is positive and would bias toward a correlation, not a null; but the gamma_90 headline claim, which supports the inside-out formation conclusion, rests on this unvalidated reuse on the incomplete-coverage subsample.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents radial stellar age gradients for 42 Local Volume dwarf galaxies, derived from spatially resolved SFHs fit to HST CMDs in four elliptical annuli per galaxy. The gradient slopes gamma_90 and gamma_50 (defined via the lookback times tau_90 and tau_50) are compared with galaxy-wide 'global' values tau_90(global) and tau_50(global), which are obtained by evaluating the same radial linear fit at the half-light radius. The central claims are that gamma_90 is significantly correlated with tau_90(global) (claimed p < 0.001), matching both the Graus et al. (2019) and Riggs et al. (2024) simulations, while gamma_50 is uncorrelated with tau_50(global), excluding the Graus et al. 'breathing-mode' prediction at the 99.98% level and favoring Riggs et al. The paper concludes that dwarf galaxies form inside-out, that age gradients are internally driven, and that the gamma_50 versus tau_50(global) relation can discriminate between feedback implementations in cosmological simulations.","tokens_in":52576,"tokens_out":8710,"duration_ms":85218,"significance":"If the central claims are robust, this is the first large, homogeneous quantitative sample of dwarf galaxy radial age gradients and it introduces a new observational discriminant between stellar feedback prescriptions in cosmological simulations. The analysis is careful in several respects: it uses Monte Carlo uncertainty propagation, tests robustness against the number of radial bins, distance and extinction assumptions, cross-checks two stellar evolution libraries (PARSEC and MIST), and validates the interpolation-based global tau values against directly measured values for 86% of the complete-coverage subset. However, the headline significance is overstated when the Spearman rank test is considered, and the interpolation-based global values are not independently validated for the 29% of the sample with incomplete azimuthal coverage, leaving the main gamma_90 correlation in need of additional verification.","major_comments":[{"comment":"The abstract and Section 4.1 claim p-values less than or equal to 0.001 for the gamma_90 versus tau_90(global) correlation, but Table 4 shows Spearman rank p-values of log10 p = -1.64 (p about 0.023) for PARSEC and log10 p = -1.03 (p about 0.09) for MIST; only the Pearson test reaches p < 0.001. Because the Spearman coefficient is reported in the same table and is less sensitive to outliers, the manuscript should either qualify the headline claim as Pearson-based or explain the discrepancy, as the current abstract overstates the significance.","section":"Abstract, §4.1, Table 4"},{"comment":"Both gamma_90 and tau_90(global) are derived from the same four-point linear fit: tau_90(global) is the best-fit line evaluated at R_hl. The validation in Section 3.3 covers only the 30 galaxies with A=1 and reports 86% agreement, leaving the 12 incomplete-coverage galaxies (29% of the sample) with no direct check. Please (a) report the covariance between gamma_90 and tau_90(global) measurement errors and propagate it in the Monte Carlo correlation analysis of Section 4.1, and (b) recompute the headline correlation using directly measured global tau_90 values for the A=1 subset (and any other galaxies for which direct estimates are possible) to demonstrate that the result is not an artifact of the interpolation on the incomplete-coverage subsample. The gamma_50 null result is less vulnerable to this issue, but the gamma_90 claim is load-bearing for the inside-out formation conclusion.","section":"§3.3, §4.1, Figs. 3 and 5"},{"comment":"The reported correlation may be sensitive to LGS3, which has both the steepest gamma_90 (8.33 Gyr/Rhl) and the earliest tau_90(global) (9.10 Gyr) and is excluded from the plotted range in Fig. 3. Please provide the correlation coefficients and p-values with LGS3 removed, and identify other high-leverage points, to confirm that the p < 0.001 result is not driven by a single galaxy.","section":"§4.1, Fig. 3, Table 3"}],"minor_comments":[{"comment":"The 'directly measured' global tau values used for the 86% validation are not defined; please state explicitly how they were computed (for example, a single SFH fit to all stars in all radial bins).","section":"§3.3"},{"comment":"The column header 'Abdundances' should read 'Abundances'.","section":"Table 1"},{"comment":"The caption does not specify whether the plotted intervals are the 16th-84th percentiles of the Monte Carlo distributions; please state this for reproducibility.","section":"Fig. 7 caption"},{"comment":"The caveat paragraph correctly notes the different definitions of tau_global between Graus et al. (present-day stellar mass) and Riggs et al. (lifetime cumulative mass), but a brief statement on the expected sign or magnitude of this methodological difference on the quantitative comparison in Fig. 7 would strengthen the discussion.","section":"§4.4"}],"recommendation":"major_revision","confidential_remarks":"The Spearman p-value overstatement and the LGS3 leverage test are concrete, easily addressable issues. The deeper risk is the interpolation-based global tau values on the incomplete-coverage subsample; requesting the covariance analysis and the A=1 direct-measurement check is appropriate before the headline correlation can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my read of 2608.07671. The genuinely new thing is the sample: 42 Local Volume dwarfs with self-consistent radial age gradient measurements (gamma_90 and gamma_50) plus global tau values, all homogeneously processed with MATCH/DOLPHOT and tested against two stellar libraries, distance/extinction priors, and bin-count variations. That catalog alone is worth having. The paper also uses the gamma_50 versus tau_50(global) non-correlation to discriminate between Graus et al. (2019) and Riggs et al. (2024), arguing against the \"breathing mode\" feedback variant. That part is the most convincing, because the null result is robust to the circularity concern: any induced covariance would bias toward a correlation, not away.\n\nThe soft spot is exactly what the stress-test flags. Both gamma_90 and tau_90(global) are derived from the same four-point linear fit per galaxy. The paper validates the interpolation at R_hl for 86% of the complete-coverage subset (30/42), but for the 12 galaxies with A<1 there is no direct check, and both axes come from the same fitted line. If incomplete azimuthal coverage biases the fit coherently, the reported p<0.001 correlation could be inflated. The authors check that coverage fraction A does not correlate with gamma, but that doesn't rule out a bias that tracks the true tau_90(global). This is a real limitation; it doesn't sink the paper, but it should be quantified—for example, with a Monte Carlo that injects azimuthal coverage patterns into the SFH fits, or a re-analysis restricted to the A=1 subset to see if the correlation persists.\n\nThe gamma_50 null result is less vulnerable, and the simulation comparison is the strongest part of the paper. The sample is archival and not volume-limited, as the authors acknowledge. The correlation with stellar mass is moderate and secondary, consistent with earlier work.\n\nOverall, this is a serious contribution. The catalog will be cited. The central inside-out claim rests partly on the gamma_90 correlation, which has this unvalidated reuse for a third of the sample. That warrants a requested revision rather than rejection. Send it to a referee who cares about SFH systematics, not just the sample size.","headline":"A substantial new catalog and a robust gamma-50 simulation discriminator, with the gamma-90 correlation weakened by a partial circularity that needs quantifying.","tokens_in":53062,"tokens_out":2259,"would_cite":true,"duration_ms":24583,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Dwarf galaxies assemble from the inside out: the radial gradient of the time to 90% of stellar mass tracks each galaxy's global star-formation history (p < 0.001), while the gradient at 50% does not, a pairing that discriminates between…","keywords":["radial stellar age gradients","dwarf galaxies","star formation histories","resolved stellar populations","color-magnitude diagram fitting","inside-out formation","stellar feedback","cosmological simulations"],"falsifier":"A decisive test: for the 12 galaxies with incomplete azimuthal coverage (areal fraction $A < 1$ in Table 1), fit SFHs to their full photometry without radial binning to obtain independent $\\tau_{90}(\\mathrm{global})$ and $\\tau_{50}(\\mathrm{global})$ values, then recompute the correlation coefficients. If the $\\gamma_{90}$–$\\tau_{90}(\\mathrm{global})$ correlation drops below $p < 0.001$, the headline signal is partly an artifact of sharing the same fitted line between both axes; if it survives, inside-out assembly is confirmed. A complementary check for the $\\gamma_{50}$ result: recompute the Graus et al. simulation predictions using the Riggs et al. definitions of global $\\tau$ (entire galaxy, all-time cumulative stellar mass) to test whether the excluded correlation is produced by the 'breathing-mode' physics or by their different radial cut at 10% of the virial radius.","tokens_in":52059,"feed_emoji":"🌌","tokens_out":15904,"duration_ms":125874,"temperature":0.7,"pith_summary":"This paper measures, for the first time on a sample of 42 galaxies, how stellar age varies with radius inside dwarf galaxies, and asks what sets those gradients. It finds that the radial gradient of $\\tau_{90}$ — the lookback time by which 90% of a dwarf's stellar mass had formed — is strongly correlated ($p < 0.001$) with the galaxy-wide value $\\tau_{90}(\\mathrm{global})$, exactly as two independent cosmological zoom-in simulations predict. The gradient of $\\tau_{50}$ shows no correlation with $\\tau_{50}(\\mathrm{global})$, a null result that agrees with one simulation family and excludes the other's 'breathing-mode' prediction at the 99.98% level. If the paper is right, dwarf galaxies form inside-out like larger galaxies, their age gradients are set internally rather than by environment, and the correlation strength between radial and global star-formation timing becomes a practical observational test for choosing among stellar-feedback implementations.","feed_headline":"Dwarf galaxy age gradients rule out 'breathing-mode' feedback","feed_subtitle":"A 42-galaxy survey shows age gradients track a dwarf's own assembly history, settling a stellar-feedback dispute.","key_machinery":"The load-bearing machinery is a uniform four-step pipeline applied to all 42 galaxies. Archival HST imaging is divided into four equally populated elliptical annuli, with shapes from 3.6 $\\mu$m surface photometry and an outer boundary at 4.4 disk scalelengths $h_r$. A star-formation history is fit to each annulus' color-magnitude diagram with the MATCH code, using both PARSEC and MIST stellar libraries; each annulus' cumulative star-formation history (CSFH) is then interpolated to give $\\tau_{90}$ and $\\tau_{50}$, the lookback times by which 90% and 50% of that annulus' stellar mass formed. A linear fit of $\\tau$ versus radius, normalized to $R_{\\rm hl}$, yields the gradient slopes $\\gamma_{90}$ and $\\gamma_{50}$, and the galaxy-wide values $\\tau_{90}(\\mathrm{global})$ and $\\tau_{50}(\\mathrm{global})$ are read off the same fitted lines at $R_{\\rm hl}$ following Graus et al. (2019). The discriminating statistic is the comparison — via Pearson and Spearman coefficients with Monte Carlo uncertainties — between the observed $\\gamma_{90}$–$\\tau_{90}(\\mathrm{global})$ and $\\gamma_{50}$–$\\tau_{50}(\\mathrm{global})$ correlations and those predicted by the two simulation families, whose differing stellar-feedback implementations produce a predicted difference specifically in the $\\gamma_{50}$ correlation.","core_discovery":"The paper's central claim is that present-day radial stellar age gradients in dwarf galaxies are a readout of each galaxy's own lifetime mass-assembly history, with measurable power to discriminate between cosmological simulations. Concretely: $\\gamma_{90}$, the slope of $\\tau_{90}$ versus radius normalized to the half-light radius $R_{\\rm hl}$, is significantly correlated ($p \\lesssim 0.001$ under both PARSEC and MIST stellar models) with $\\tau_{90}(\\mathrm{global})$, in agreement with both the Graus et al. (2019) and Riggs et al. (2024) simulations; while $\\gamma_{50}$ is uncorrelated with $\\tau_{50}(\\mathrm{global})$, excluding the Graus et al. prediction at the 99.98% level and favoring Riggs et al. The simulation difference being adjudicated is the 'breathing mode' of feedback-driven radial outflows of star-forming gas, which imprints an initial radial velocity on young stars in the Graus et al. simulations but is absent in the Riggs et al. simulations. The authors further report that no target galaxy shows a significantly inside-out present-day gradient, and that no environmental metric (tidal index $\\theta_1$, nearest-neighbor distance $D(\\mathrm{NN})$, or distance to a $10^{10}\\,M_\\odot$ host) correlates with either gradient. They conclude that dwarfs form inside-out, that the observed outside-in-to-flat range of gradients results from feedback-driven outward migration of older stars combined with growing birth radii of young populations, and that radial age gradients are an actionable observational discriminant between simulations with different stellar feedback implementations.","pith_inferences":["Because $\\tau_{90}(\\mathrm{global})$ and $\\tau_{50}(\\mathrm{global})$ are interpolated from the same four-point radial fits that define $\\gamma_{90}$ and $\\gamma_{50}$, the two axes of the headline correlation are not statistically independent; recomputing global $\\tau$ values from SFHs fit to complete galaxy-wide CMDs for the 12 targets with partial azimuthal coverage would test whether the $p < ","The same annulus-and-CSFH pipeline extends naturally to lower masses: the Riggs et al. prediction that reionization-quenched dwarfs below $\\mathrm{Log}\\,M_\\star \\sim 6$ have flat gradients is directly testable with deep JWST imaging of the least massive Local Group dwarfs.","A second decisive experiment suggested by the paper's own caveats: re-deriving the Graus et al. predictions with Riggs et al.'s definitions of global $\\tau$ (full galaxy, all-time cumulative mass) would show whether the excluded $\\gamma_{50}$ correlation is a product of the feedback model itself or of the different radial and definitional cuts the two simulations used.","The null $\\gamma_{50}$ result has a physical reading the paper leaves implicit: the radial ordering of the more recently formed half of a dwarf's stars is governed by stochastic, local processes rather than a coherent galaxy-wide timing relation, a distinction next-generation simulations could isolate directly."],"forward_implications":["Dwarf galaxies assemble their stars inside-out like more massive disk galaxies, with feedback-driven outward migration of old stars slightly outpacing the outward creep of recent star formation.","None of the 42 dwarfs shows a statistically significant inside-out present-day gradient, setting a quantitative target for simulations: gradients should range from outside-in to flat, not inverted.","Environment does not set internal age gradients: tidal index, nearest-neighbor distance, and distance to a $\\geq 10^{10}\\,M_\\odot$ host all fail to correlate with $\\gamma_{90}$ or $\\gamma_{50}$, implying internal drivers dominate in this mass range.","The null $\\gamma_{50}$–$\\tau_{50}(\\mathrm{global})$ correlation is a new quantitative constraint on stellar-feedback physics, arguing against feedback that simultaneously drives radial breathing motions of star-forming gas.","Gradient-versus-global-SFH correlation strength is an actionable parameter: any future simulation's feedback implementation can be ranked by whether it reproduces the observed $\\gamma_{90}$ correlation and $\\gamma_{50}$ non-correlation together."],"supporting_citations":[{"why":"Supplies the predicted $\\gamma_{90}$–$\\tau_{90}(\\mathrm{global})$ and $\\gamma_{50}$–$\\tau_{50}(\\mathrm{global})$ correlations from FIRE-2 dwarfs, plus the $R_{\\rm hl}$-interpolation method for global $\\tau$ values that the paper adopts.","marker":"A. S. Graus et al. (2019)"},{"why":"Provides the competing simulation set predicting no $\\gamma_{50}$–$\\tau_{50}(\\mathrm{global})$ correlation, which the observations favor, and the mass-selected comparison subsample used in Figures 5–7.","marker":"C. L. Riggs et al. (2024)"},{"why":"The FIRE-2 simulation physics underlying the Graus et al. predictions, including the feedback implementation that produces 'breathing modes'.","marker":"P. F. Hopkins et al. (2018)"},{"why":"Identifies the breathing-mode mechanism — feedback-driven radial migration and population gradients — that the $\\gamma_{50}$ null result is used to adjudicate.","marker":"K. El-Badry et al. (2016)"},{"why":"The MATCH code that fits per-annulus star formation histories, the computational core of every $\\tau_{90}$ and $\\tau_{50}$ measurement.","marker":"A. E. Dolphin (2002)"},{"why":"Prescribes the random uncertainties on the SFHs that propagate into the asymmetric $\\tau$ uncertainties and hence the gradient fits.","marker":"A. E. Dolphin (2013)"},{"why":"Provides the 3.6 $\\mu$m ellipse fits (center, position angle, axis ratio, scalelengths, half-light radii) that define the four radial bins for most targets.","marker":"McQuinn et al. (in prep.)"},{"why":"The prior quantitative gradient measurement for WLM that validates the pipeline against independent JWST-based SFHs and supplies a comparison value.","marker":"R. E. Cohen et al. (2025)"}],"fun_headline_variants":["Dwarf age gradients pick simulation winner","42 dwarfs: age gradients expose feedback physics","Dwarf galaxy ages rule out breathing-mode feedback","Age gradients in 42 dwarfs trace own assembly","Stellar age gradients settle dwarf feedback dispute"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that reading a galaxy's global star-formation timing off the same radial line that defines its gradient — evaluated at one half-light radius — is unbiased; the shortcut was checked only on galaxies with complete sky coverage, where it agreed within uncertainties in 86% of cases, so the 29% of the sample with incomplete azimuthal coverage is precisely where a bias would leak into both axes of the headline correlation.","fun_headline_variants_meta":{"raw":{"variants":["Dwarf age gradients pick simulation winner","42 dwarfs: age gradients expose feedback physics","Dwarf galaxy ages rule out breathing-mode feedback","Age gradients in 42 dwarfs trace own assembly","Stellar age gradients settle dwarf feedback dispute"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1467,"prompt_tokens":1183,"completion_tokens":284,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":799,"completion_tokens_details":{"reasoning_tokens":215}},"tokens_in":799,"tokens_out":284,"duration_ms":3793,"temperature":1.0,"reasoning_tokens":215,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:25:25.940305+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test: for the 12 galaxies with incomplete azimuthal coverage (areal fraction $A < 1$ in Table 1), fit SFHs to their full photometry without radial binning to obtain independent $\\tau_{90}(\\mathrm{global})$ and $\\tau_{50}(\\mathrm{global})$ values, then recompute the correlation coefficients. If the $\\gamma_{90}$–$\\tau_{90}(\\mathrm{global})$ correlation drops below $p < 0.001$, the headline signal is partly an artifact of sharing the same fitted line between both axes; if it survives, inside-out assembly is confirmed. A complementary check for the $\\gamma_{50}$ result: recompute the Graus et al. simulation predictions using the Riggs et al. definitions of global $\\tau$ (entire galaxy, all-time cumulative stellar mass) to test whether the excluded correlation is produced by the 'breathing-mode' physics or by their different radial cut at 10% of the virial radius.","supporting_citations":[],"review_version":1}