{"id":"b45ba275-ae6e-4da5-9526-91f69ac4889e","arxiv_id":"2411.17882","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"At z~0.3, simulated galaxies match the observed global star-forming main sequence but show shallower resolved star-forming main sequence slopes and different radial star formation profiles from MAGPI, with central suppression varying between simulations due to AGN feedback prescriptions.","lead":"This study compares the internal distribution of star formation in galaxies at z~0.3, using MAGPI observations and three large cosmological simulations (EAGLE, Magneticum, IllustrisTNG). It finds that global star formation scaling relations agree, but radial profiles within galaxies often do not, with differences in central suppression linked to how each simulation implements black hole feedback.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Resolved SFMS is estimated with different methods for MAGPI (H-alpha spaxel fit) and simulations (PDF peak over all non-zero SFR spaxels), so the claimed slope discrepancy and the ΔΣSFR profiles that depend on it may be method artifacts rather than astrophysics.","rationale":"The reader's weakest assumption is exactly the one I find most load-bearing: the resolved SFMS is not measured apples-to-apples between MAGPI and the simulations, yet it anchors both the slope-discrepancy claim and the ΔΣSFR normalization used for radial profiles. The paper is otherwise careful and transparent—SimSpin mock cubes, PSF/pixel matching, stellar/halo mass cuts, and the appendix showing percentile spreads are all to its credit. The concern is not that the authors are wrong, but that the central quantitative claim is currently underdetermined by the analysis as presented. A single re-analysis applying a common estimator to both datasets would settle it. The reader's CONDITIONAL verdict is appropriate; I am not moving it because the qualitative conclusions (e.g., central suppression differences among simulations) have independent support from the simulation-internal comparisons and from prior literature, and the paper already flags the error caveat in Section 4.5. However, if the common-estimator test showed method invariance, the resolved-SFMS claim would be strengthened; if not, the headline would need to be weakened from a quantitative slope disagreement to a qualitative tracer/selection-dependent difference. Hence UNCHANGED conditional acceptance.","tokens_in":31436,"tokens_out":7154,"duration_ms":64899,"concrete_test":"Recompute the MAGPI resolved SFMS using the exact simulation estimator of §4.4: bin MAGPI Hα star-forming spaxels by Σ★, estimate the ΣSFR PDF with Silverman-rule Gaussian KDEs, take the peak of each bin, and fit peaks with ODR. If the slope becomes consistent with the simulation slopes (~0.73–0.81) rather than 0.92±0.01, the headline resolved-SFMS discrepancy is a method artifact. Then recompute Fig. 4 ΔΣSFR profiles with this common estimator to check whether radial-profile disagreements persist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's first quantitative result—'the slope of the resolved SFMS does not agree within 1–2σ' (abstract; Table 2)—and the radial-profile comparison built on ΔΣSFR (Eq. 2) rest on the comparability of two different resolved-SFMS estimators. MAGPI: direct fit to Hα-detected, BPT-classified star-forming spaxels with dust-corrected Hα SFRs (Paper I; §4.4). Simulations: peak of a Gaussian-kernel ΣSFR PDF in Σ★ bins, fitted with ODR, using all spaxels with non-zero SFR after a uniform detection floor (§4.4; Fig. 3). These differ in spaxel selection (MAGPI excludes AGN/composite and non-detections; simulations include all star-forming gas), SFR tracer (Hα vs instantaneous), and fitting statistic (spaxel values vs PDF mode). The observed offset (MAGPI 0.92±0.01; EAGLE 0.80±0.02, Magneticum 0.73±0.07, TNG 0.81±0.02) is in the direction expected if the simulation PDF peak is dragged down at high Σ★ by low-sSFR spaxels that MAGPI would exclude. Because ΔΣSFR is defined relative to each sample's own resolved SFMS, the radial profiles in Fig. 4 and the 'only far below the SFMS' agreement are also sensitive to this mismatch. The paper acknowledges the different approaches (§4.4) but never validates them against one another, so the central claim that resolved profiles reveal discrepancies hidden by global relations is not yet secure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares spatially resolved star formation at z~0.3 between MAGPI observations and mock MUSE observations of galaxies drawn from the EAGLE, Magneticum, and IllustrisTNG cosmological simulations, using SimSpin to match PSF, pixel scale, and spaxel-by-spaxel SFR and stellar-mass detection limits. It reports that the global star-forming main sequence (SFMS) slopes agree within 1-2 sigma, while the resolved SFMS slopes are shallower in all three simulations than in MAGPI (Table 2). It then constructs DeltaSigma_SFR radial profiles for galaxies in different DeltaSFR bins, finding that simulations and MAGPI agree only for galaxies far below the SFMS, that central suppression within R~1.5 Re differs among simulations, and that central versus satellite galaxies show distinct environmental trends.","tokens_in":31775,"tokens_out":6067,"duration_ms":59662,"significance":"If the resolved-SFMS discrepancy and the radial-profile differences are astrophysical rather than methodological, the paper would demonstrate that spatially resolved star formation provides a discriminating test of subgrid feedback models beyond global scaling relations. The study is timely and the observational matching effort is substantial: the authors use mock observations, match detection limits and parameter ranges, and analyze three independent simulation codes. The central result, however, hinges on the comparability of two different resolved-SFMS estimators, and the current manuscript does not validate that comparability, so the quantitative conclusions are not yet secure.","major_comments":[{"comment":"The load-bearing claim that the resolved SFMS slope does not agree within 1-2 sigma between MAGPI and the simulations rests on comparing two different estimators. MAGPI's resolved SFMS is derived from a direct fit to H-alpha-detected, BPT-classified star-forming spaxels, whereas the simulations use the peak of the Sigma_SFR probability density over all non-zero SFR spaxels, fitted with ODR. These differ in spaxel selection, SFR tracer, and fitting statistic. The simulation PDF peak may be systematically dragged down at high Sigma_star by low-sSFR spaxels that MAGPI would exclude, which could explain the shallower slopes (MAGPI 0.92+-0.01 versus EAGLE 0.80+-0.02, Magneticum 0.73+-0.07, IllustrisTNG 0.81+-0.02). Because DeltaSigma_SFR in Eq. (2) is measured relative to each sample's own resolved SFMS, the radial-profile comparison in Fig. 4 and the claim of agreement only far below the SFMS are also affected. The authors acknowledge the different approaches in Section 4.4 but never validate them against one another; applying the same resolved-SFMS estimator to both datasets, or otherwise demonstrating that the slope difference survives, is required.","section":"Section 4.4, Table 2, Fig. 3"},{"comment":"The quoted inner and outer profile slopes and their errors are bootstrap standard errors on medians. With thousands of simulated galaxies and many spaxels per bin, these errors are extremely small (often <0.01 dex/Re), while the galaxy-to-galaxy scatter is roughly 0.4 dex, as the authors note in Section 4.5 and illustrate in Appendix B. The abstract's 'does not agree within 1-2 sigma' statement and the interpretation of slope differences in Section 5.2 therefore rely on error bars that understate the true population scatter. The authors should report the scatter, use a mixed-effects or hierarchical model, or otherwise present a significance measure that reflects the galaxy-to-galaxy variance. Additionally, the choice of 1.5 Re as the inner/outer division is an ad hoc assumption (Section 5.1); a sensitivity test using other cutoffs would strengthen the slope comparisons in Table 3.","section":"Section 5, Table 3, Fig. 4"},{"comment":"The abstract attributes the differences in central suppression within R~1.5 Re to different AGN feedback prescriptions. The three simulations differ simultaneously in hydrodynamic scheme, resolution, stellar feedback implementation, BH seeding, and calibration targets, so the comparison is not a controlled experiment. The interpretation is plausible and consistent with previous literature, but as stated it overreaches. The authors should either temper the causal attribution or support it with an analysis that controls for at least some of the other differences, such as comparing feedback variants within the same simulation code or explicitly discussing how resolution and seeding affect the central profiles.","section":"Abstract and Section 6.1"}],"minor_comments":[{"comment":"The definitions of log10(SFR_MS) and log10(Sigma_SFR,MS) in Eqs. (1) and (2) are not written explicitly as functions of stellar mass; making the functional dependence clear would help readers reproduce the DeltaSFR and DeltaSigma_SFR calculations.","section":"Section 4.4, Eq. (1)"},{"comment":"The caption refers to 'turquoise data points' for the simulation SFMS fits, but the figure appears to use a different color scheme; the caption should be aligned with the actual plot.","section":"Fig. 3 caption"},{"comment":"There are several grammatical and word-choice issues, such as 'the latter for which may not be reflected in use of SFR indicators' in Section 1 and the use of 'i.e.' where 'e.g.' is meant in Sections 4.2 and 5.1. A careful language edit would improve clarity.","section":"Throughout"},{"comment":"The statement that bootstrap errors 'may not represent the overall scatter' is important but could be emphasized more strongly; the main text frequently discusses 'discrepancies' without recalling this caveat.","section":"Section 4.5"}],"recommendation":"major_revision","confidential_remarks":"This is a solid observational-theoretical comparison with a careful mock-observation pipeline. The main concern is the unmatched resolved-SFMS estimators, which is fixable with additional analysis and would determine whether the headline result survives. I would not reject the paper, but I would require the validation before publication. The attribution to AGN feedback is also stronger than the evidence supports, though this can be tempered with language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things before reading. First, this is the closest thing to an apples-to-apples resolved star formation comparison at z~0.3 I have seen: the authors run SimSpin on EAGLE, Magneticum, and IllustrisTNG, match SFR and stellar mass detection limits, PSF, pixel scale, and then apply the same radial profile machinery to both sides. Second, the headline quantitative claim—resolved SFMS slopes disagree between MAGPI and all three simulations—is built on two different estimators that are never cross-checked. That discrepancy is real in the paper, but it may be a method artifact rather than astrophysics.\n\nWhat is genuinely new: the consistent mock-observation pipeline applied to three simulations with central/satellite and halo-mass splits. The authors show that environmental trends in the profiles only appear when controlling for both central/satellite status and halo mass, which is a nice point and rings true. They also show central suppression within ~1.5 Re differs across the simulations in a way that tracks their known AGN feedback implementations—EAGLE weak in the centre, Magneticum and TNG stronger—consistent with prior literature. That simulation-to-simulation comparison uses identical methods on both sides, so it is robust to the MAGPI estimator issue. The paper is honest about non-detections, D4000 upper limits, preliminary halo masses, and the fact that bootstrap errors on medians hide ~0.4 dex galaxy-to-galaxy scatter (Appendix B shows the percentiles).\n\nSoft spots, in order of severity. First, the resolved SFMS is measured for MAGPI from H-alpha/BPT-selected spaxels but for simulations from the peak of the Sigma_SFR PDF over all non-zero SFR spaxels. The observed slope offset (0.92 vs ~0.8) is in the direction you would expect if the simulation PDFs are dragged down at high Sigma_star by low-sSFR spaxels that MAGPI would classify as non-star-forming. The authors acknowledge the different approaches in Section 4.4 but never apply the same method to both datasets. That is the load-bearing weakness. Second, the Table 3 slope errors are standard errors on medians and are often <0.01 dex/Re; the paper uses 1–2 sigma language throughout, which overstates significance even though it flags the scatter later. Third, the MAGPI sample is not environment-matched to the simulations—it is group-centric and projected within ~270 kpc—so the satellite comparison is approximate. The authors say this plainly. Fourth, reproducibility depends on future data releases.\n\nNone of this kills the paper. The main qualitative result—global relations hide discrepancies that radial profiles reveal, and AGN feedback implementation shapes central suppression—is supported by the figures and by prior work. The specific resolved SFMS slope mismatch should be re-measured with a common estimator before being cited as quantitative evidence.\n\nFor whom: anyone working on resolved star formation, IFU surveys, or simulation calibration. It deserves a serious referee, but the referee should push for the estimator validation and a scatter-honest error treatment. I would not cite the resolved SFMS slope as it stands, but I would cite the central/satellite and halo-mass profile comparison.","headline":"A careful mock-observation comparison of MAGPI radial star formation with three simulations whose qualitative results are worth engaging, but the headline resolved-SFMS slope mismatch rests on two unmatched estimators and needs a validation test before it is cited as quantitative evidence.","tokens_in":32465,"tokens_out":1854,"would_cite":true,"duration_ms":20957,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"At $z\\sim0.3$, spaxel-resolved star-formation profiles in MAGPI disagree with all three cosmological simulations even though global main-sequence slopes agree, and the central suppression difference tracks active-galactic-nucleus feedback…","keywords":["galaxy evolution","star formation","integral field spectroscopy","cosmological simulations","star-forming main sequence","AGN feedback","galaxy environment","radial profiles"],"falsifier":"Re-fit the resolved star-forming main sequence for MAGPI and every simulation using one common method—for instance, applying the simulation PDF-peak fit to MAGPI spaxels or applying the MAGPI linear fit to simulation spaxels. If the slopes then agree within 1$\\sigma$, the reported resolved-SFMS disagreement is a fitting artifact; if they still disagree, the claim of a physical discrepancy survives.","tokens_in":31191,"feed_emoji":"🔭","tokens_out":11024,"duration_ms":86465,"temperature":0.7,"pith_summary":"This paper tries to show that how star formation is arranged inside a galaxy—not just how much star formation it has in total—is a sharper test of cosmological simulations. Using MAGPI integral-field observations at $z\\sim 0.3$ and mock observations of EAGLE, Magneticum, and IllustrisTNG built with SimSpin, the authors find that the slope of the resolved (per-spaxel) star-forming main sequence disagrees between MAGPI and all three simulations at the $1$–$2\\sigma$ level, while the global star-forming main sequence agrees. The paper argues that the spatial pattern of disagreement, particularly central star-formation suppression within $\\sim1.5\\,R_e$, tracks how each simulation implements active galactic nucleus feedback. If true, radial star-formation profiles become a discriminating probe of subgrid physics that global scaling relations cannot provide.","feed_headline":"Radial star-formation maps disagree with three galaxy simulations","feed_subtitle":"Global main-sequence slopes agree, but spaxel-by-spaxel profiles reveal feedback differences that galaxy-wide averages hide.","key_machinery":"The load-bearing object is the resolved star-forming main sequence and the radial offset from it, $\\Delta\\Sigma_{\\rm SFR}=\\log_{10}(\\Sigma_{\\rm SFR,spax})-\\log_{10}(\\Sigma_{\\rm SFR,MS})$, measured in elliptical annuli of width $0.5\\,R_e$ out to $5\\,R_e$. The resolved SFMS is the per-spaxel relation between star-formation-rate surface density and stellar-mass surface density; MAGPI's version is built from dust-corrected H$\\alpha$ star-forming spaxels, while each simulation's version is built from the peak of the $\\Sigma_{\\rm SFR}$ probability density in bins of $\\Sigma_*$. SimSpin mock data cubes impose MAGPI's PSF, line-spread function, pixel scale, and SFR/stellar-mass surface-density detection limits on the simulations, so the comparison is meant to be an instrument-matched one.","core_discovery":"The central claim is that spaxel-resolved star-formation profiles at $z\\sim0.3$ separate the MAGPI observations from all three simulations in a way that global measurements hide. The resolved star-forming main sequence fitted to MAGPI's H$\\alpha$-detected star-forming spaxels has a steeper slope ($0.92\\pm0.01$) than the slopes obtained from the peak of the $\\Sigma_{\\rm SFR}$ probability distribution in each simulation ($0.73$–$0.81$), a disagreement outside $1$–$2\\sigma$. In radial $\\Delta\\Sigma_{\\rm SFR}$ profiles, the simulations only match the observed inside-out quenching signature for galaxies far below the main sequence; for galaxies on or just below it, the simulations show differing central suppression within $\\sim1.5\\,R_e$, which the paper attributes to different AGN feedback prescriptions (single-mode thermal feedback in EAGLE versus dual-mode thermal/kinetic feedback in Magneticum and IllustrisTNG). The paper further claims that centrals and satellites follow different radial quenching paths, with centrals showing halo-mass-dependent central suppression and satellites showing increasing outskirts suppression, and that these environmental trends only appear when both central/satellite status and halo mass are controlled.","pith_inferences":["Inference: The resolved-SFMS slope gap might partly reflect the different SFR tracers, since H$\\alpha$ traces roughly 10 Myr of star formation while simulations report instantaneous SFRs; averaging simulated SFRs over about 10 Myr before fitting would test whether the slope disagreement is physical.","Inference: Applying the same fitting algorithm, either the PDF-peak method or the direct linear fit, to both MAGPI and simulation spaxels would isolate whether the reported resolved-SFMS disagreement is a method artifact, a test the paper does not perform.","Inference: If the central-suppression attribution to AGN feedback is correct, simulations that toggle between kinetic and thermal AGN modes at fixed resolution should reproduce the Magneticum/IllustrisTNG versus EAGLE ordering in central slopes, offering a clean falsification test.","Inference: The environmental result implies that group-scale integral-field surveys need sample sizes large enough to bin by both central/satellite status and halo mass, which may push future wide-field spectrographs to prioritize depth over field of view."],"forward_implications":["Resolved star-formation scaling relations are a model discriminator even when the global star-forming main sequence is reproduced within $1$–$2\\sigma$.","Galaxies far below the star-forming main sequence show inside-out quenching in both MAGPI and all three simulations, so this quenching mode is robust across feedback implementations.","Differences in central suppression within $\\sim1.5\\,R_e$ can be used to distinguish AGN feedback prescriptions, with the strongest suppression appearing in simulations that inject kinetic or dual-mode AGN feedback.","Environmental quenching is visible in radial profiles only when galaxies are split by central/satellite status and halo mass; population-averaged profiles wash it out.","Mock observations that match PSF, pixel scale, and detection limits are necessary for any such comparison, because resolution and selection effects change the measured radial trends."],"supporting_citations":[{"why":"Paper I: supplies the MAGPI observed radial profiles, resolved SFMS definition, and $\\Delta$SFR galaxy classifications that this paper compares against.","marker":"Mun et al. (2024)"},{"why":"Provides SimSpin, the code that builds the mock data cubes with MAGPI PSF, pixel scale, and line-spread function.","marker":"Harborne et al. 2020, 2023"},{"why":"Defines the EAGLE simulation and its stochastic thermal AGN feedback model, the comparison baseline for central suppression.","marker":"Schaye et al. 2015"},{"why":"Defines the Magneticum simulation and its dual-mode AGN feedback prescription, used for the central suppression comparison.","marker":"Teklu et al. 2015"},{"why":"One of the IllustrisTNG model papers underlying the TNG100-1 run and its kinetic/thermal AGN feedback implementation.","marker":"Nelson et al. 2018"},{"why":"Supplies the $\\Delta$SFR bin definitions (above, on, just below, far below the SFMS) that organize all radial-profile comparisons.","marker":"Bluck et al. 2020"},{"why":"Introduces the resolved SFMS and $\\Delta\\Sigma_{\\rm SFR}$ offset method that this paper adapts to MAGPI and simulations.","marker":"Ellison et al. 2018"},{"why":"Provides the $z\\sim1$ precedent where TNG50 with kinetic AGN feedback matches observed inside-out quenching, the benchmark this paper's conclusions extend to $z\\sim0.3$.","marker":"Nelson et al. 2021"},{"why":"Documents the central star-formation excess and young central stellar populations in EAGLE galaxies, used to interpret EAGLE's lack of central suppression.","marker":"Lagos et al. 2022"}],"fun_headline_variants":["Spaxel-resolved star formation disagrees with simulations","MAGPI: small-scale star formation breaks simulation consensus","Radial star formation: observations vs. three simulations","Simulations miss MAGPI's resolved star-forming main sequence","Feedback models split radial star formation in simulations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the resolved star-forming main sequence measured from MAGPI's dust-corrected H$\\alpha$ star-forming spaxels and the one measured from each simulation's peak of the $\\Sigma_{\\rm SFR}$ probability density over all non-zero SFR spaxels are the same quantity; the paper never applies a single fitting method to both datasets.","fun_headline_variants_meta":{"raw":{"variants":["Spaxel-resolved star formation disagrees with simulations","MAGPI: small-scale star formation breaks simulation consensus","Radial star formation: observations vs. three simulations","Simulations miss MAGPI's resolved star-forming main sequence","Feedback models split radial star formation in simulations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000881,"raw_usage":{"total_tokens":3885,"prompt_tokens":1104,"completion_tokens":2781,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":2705}},"tokens_in":720,"tokens_out":2781,"duration_ms":21400,"temperature":1.0,"reasoning_tokens":2705,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:44:09.372657+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-fit the resolved star-forming main sequence for MAGPI and every simulation using one common method—for instance, applying the simulation PDF-peak fit to MAGPI spaxels or applying the MAGPI linear fit to simulation spaxels. If the slopes then agree within 1$\\sigma$, the reported resolved-SFMS disagreement is a fitting artifact; if they still disagree, the claim of a physical discrepancy survives.","supporting_citations":[],"review_version":1}