{"id":"95dd2004-ad91-4d9b-a32d-e0dc409aacdf","arxiv_id":"2603.01726","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"EAGLE and TNG100 both reproduce the observed trend that galaxies at fixed stellar mass are more compact when their star formation deviates from the main sequence, though EAGLE predicts a stronger and TNG100 a weaker dependence than observed.","lead":"This paper compares the joint stellar-mass–star-formation-rate–size relation of galaxies in two cosmological simulations (EAGLE and TNG100) with observational samples from SDSS and CANDELS. It finds that both simulations reproduce the observed trend that galaxies become more compact as they move away from the star-forming main sequence, but with different strengths.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Untested half-mass-to-half-light proxy may bias the reported slope difference between simulations and observations, weakening the central claim of differential trend strength.","rationale":"The reader's conditional verdict centers on exactly this issue: the direct adoption of 3D half-mass radius as a proxy for observed half-light radius, with no forward-modeling. I agree this is the most load-bearing concern because the paper's headline claim is about the relative strength of the size–SFMS-offset trend, not just its existence. A constant offset between R_hsm and R_e would indeed cancel in the relative-size metric, but a trend in that ratio with ΔlogSFR would not. The authors themselves acknowledge the need for forward-modeling in Section 6, and they cite previous work (Katsianis et al. 2020) showing that matched methods can change conclusions. The concrete test of generating mock observations and re-measuring sizes would directly settle whether the EAGLE/observed and TNG100/observed differences in slope are robust. I also note that the paper lacks formal uncertainty quantification on the slopes, which is a secondary concern, but the proxy issue is more fundamental because it threatens the validity of the comparison itself. Therefore the verdict should remain CONDITIONAL (i.e., UNCHANGED from the reader), pending this test.","tokens_in":20865,"tokens_out":7668,"duration_ms":78611,"concrete_test":"For EAGLE and TNG100 at 0.5<z<2.5, generate mock observations with SKIRT radiative transfer (or an equivalent empirical conversion) to produce rest-frame optical images, measure half-light radii with the same Sérsic fitting and SFR-ladder methodology as the CANDELS catalog, then recompute ΔlogRe as a function of ΔlogSFR using the same SFMS fitting. If the EAGLE-vs-observed and TNG-vs-observed differences in slope remain after this forward-modeling, the central claim survives; if the slopes converge or reverse, the reported strength comparison is an artifact of the size proxy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 explicitly states 'we do not generate synthetic images... directly adopt the cataloged half–stellar–mass radius, R_hsm, as a proxy for the observed effective radius R_e.' The relative-size metric ΔlogRe removes only constant offsets between the two size definitions; it does not remove a dependence of R_hsm/R_e on SFMS offset or redshift. For galaxies far above the SFMS, young stellar populations, dust, and profile shape (Sérsic index) can shift the light-weighted radius relative to the mass-weighted radius by amounts comparable to the trend strength (≈0.1–0.2 dex per dex in ΔlogSFR, as suggested by Fig. 4). If this ratio varies with ΔlogSFR, the EAGLE vs observed slope difference claimed in Section 5.2 could be exaggerated or even inverted. The authors acknowledge the limitation in Section 6 and cite Katsianis et al. (2020), who showed that matching SFR derivation methods can alleviate tensions, but no quantitative test is provided. Since the headline result is specifically about the strength of the trend, this untested proxy assumption is the most load-bearing gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares the joint stellar mass–SFR–size relation in two cosmological hydrodynamical simulations (EAGLE and TNG100) with observations from SDSS (z<0.2) and CANDELS (0.5<z<2.5). Using a relative size metric ΔlogRe defined from the mass–size relation of star-forming main-sequence (SFMS) galaxies, the authors report that both simulations reproduce the observed trend that galaxy sizes decrease with increasing offset from the SFMS ridge, that this trend weakens and is not detected in the observed sample at 1.5<z<2.5, and that EAGLE predicts a stronger size dependence while TNG100 predicts a weaker one. The paper interprets the trend as an emergent property of the simulations and relates it to time variability in star formation.","tokens_in":21244,"tokens_out":2098,"duration_ms":23384,"significance":"If the qualitative and differential claims are robust, the result is significant: the joint mass–SFR–size relation would provide a new, non-trivially emergent constraint on galaxy formation models, with different simulation implementations making different predictions. The paper's use of two independent simulation suites and its choice of a relative size metric to remove mass dependence are sensible, and the qualitative comparison is clearly presented. However, the central quantitative claims—especially the EAGLE-vs-observed and TNG-vs-observed differences in trend strength—rest on two unquantified assumptions: the equivalence of 3D half-mass radius to observed half-light radius in a relative sense, and the absence of significant measurement/scatter uncertainties. Because these are not tested, the paper is better viewed as presenting a promising qualitative result that needs strengthening before the quantitative conclusions can be accepted.","major_comments":[{"comment":"The analysis directly adopts the 3D stellar half-mass radius R_hsm as a proxy for the observed projected half-light radius R_e, and the paper acknowledges in Section 6 that forward-modeling is needed. The relative metric ΔlogRe removes only constant offsets between the two size definitions; it does not remove a dependence of R_hsm/R_e on ΔlogSFR or on redshift. For galaxies far above the SFMS, young stars, dust, and varying Sérsic index can shift the light-weighted radius relative to the mass-weighted radius by amounts comparable to the trend strength shown in Fig. 4 (≈0.1–0.2 dex per dex in ΔlogSFR). If this ratio varies with ΔlogSFR, the claimed EAGLE-vs-observed and TNG-vs-observed slope differences in Section 5.2 could be biased or even inverted. The limitation is acknowledged, but no quantitative test is provided. I request at least a check using the TNG100 2D half-light radii (whic","section":"Section 4 and Section 6 (R_hsm vs R_e proxy)"},{"comment":"The central claims—that the trend is 'not detected' at 1.5<z<2.5 in observations, and that EAGLE predicts a stronger and TNG100 a weaker dependence—are stated without any uncertainty quantification. Fig. 4 shows median lines and 16th–84th percentile shaded regions, but no confidence intervals on the medians, no significance tests for the slope of ΔlogRe versus ΔlogSFR, and no assessment of whether the EAGLE/observed difference is statistically meaningful. The high-z CANDELS sample has only 5,111 galaxies above the mass cut (Table 1), and the measurement uncertainties are plausibly large, but this is asserted rather than demonstrated. Please provide bootstrap/jackknife confidence intervals on the median trends and, ideally, fitted slopes with uncertainties, so the reader can judge whether the differential claims are supported.","section":"Section 5.2 and Fig. 4 (error bars and significance)"},{"comment":"The iterative SFMS fitting is described briefly, and the mass–size relation is fitted only to the SFMS-selected galaxies. The choice of a 1 dex lower threshold and the exclusion of zero-SFR galaxies could affect the resulting ΔlogSFR distribution and thus the ΔlogRe–ΔlogSFR relation. No convergence criterion, error on the fitted parameters, or test of the sensitivity to the threshold is given. This is not fatal, but since the offsets ΔlogSFR are defined relative to the fitted SFMS, a more detailed description or a robustness test would help support the quantitative comparison.","section":"Section 4 (SFMS and mass–size fitting procedure)"}],"minor_comments":[{"comment":"The caption labels the color axis as 'Δ log SFR' but the text and figure description indicate the color scale should be ΔlogRe. Please correct the caption.","section":"Figure 3 caption"},{"comment":"There are typographical issues, e.g., 'Fig.,1' in the Figure 3 caption, 'amnd' in the Fortuné et al. reference, and inconsistent use of 'Sersic' vs 'Sérsic'. These should be cleaned up.","section":"General typography"},{"comment":"The discussion of the low-mass feature ('the largest galaxies from each of the stellar mass bins collectively follow a somewhat steeper log M*–log SFR relation') is interesting but not quantified; a short explanation of what is meant by 'steeper' would improve clarity.","section":"Section 5.1"},{"comment":"The photometric-redshift uncertainties for CANDELS are not discussed; given that the SFRs and masses come from SED fitting, some statement about redshift-dependent completeness or uncertainties (beyond the qualitative sentence in Section 5.1) would be useful.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The qualitative finding is interesting and the paper is well positioned, but the two load-bearing gaps—lack of forward-modeling for the size proxy and lack of statistical uncertainty quantification—need to be addressed before the differential claim can be considered robust. The authors have access to the TNG100 2D half-light radii, which would provide a relatively low-cost check. I do not see a fundamental circularity in the analysis; the joint relation is not fitted to the simulations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThe new thing here is the direct comparison of the joint mass–SFR–size relation in two major cosmological simulations against SDSS and CANDELS, using the same relative-size metric on both sides. That specific comparison has not been published before. The main qualitative result—galaxies get more compact away from the SFMS ridge, in both sims and observations—is visually robust, and it is genuinely emergent: neither EAGLE nor TNG100 was calibrated to this joint relation. The quantitative wrinkle, EAGLE predicting a stronger size dependence on SFMS offset than observed and TNG100 a weaker one, is a useful handle on feedback physics if it holds.\n\nThe paper does several things well. The relative-size definition removes most of the mass dependence and makes the sim-observation comparison cleaner than comparing absolute radii. The iterative SFMS fitting is applied consistently across datasets. They check central vs total samples in SDSS and find no important difference. The interpretation in terms of SFR time variability is clearly argued and properly seeded in earlier work. They also state the main limitation in Section 6 rather than burying it.\n\nNow the soft spots, in proportion. The biggest one is exactly the stress-test: using 3D half-mass radius as a proxy for observed half-light radius. The relative metric removes constant offsets, but not a dependence of R_hsm/R_e on SFMS offset or redshift. For starbursts there are plausible changes in Sérsic index, dust, and young light weighting that shift the light radius relative to the mass radius by amounts comparable to the 0.1–0.2 dex trend. The authors acknowledge this and cite Katsianis et al. (2020), but they do not quantify it. Because the headline is specifically about the strength of the trend, that gap is load-bearing. I would not call it fatal—a mock run with SKIRT or similar could settle it—but as written the differential claim is conditional.\n\nSecond, there are no error bars or significance tests anywhere. The high-z observational non-detection is explained by 'increased measurement uncertainties' without a number. The 16th–84th percentile shading helps but does not replace bootstrap or fitting uncertainties. Third, the EAGLE and TNG size definitions differ (30 kpc aperture vs total subhalo) and the SFMS normalizations are offset from observations at high z; they use each sample's own SFMS, which is the right call, but those offsets limit how clean the ΔSFR comparison is.\n\nWho's this for? People working on galaxy formation simulations and on using size–SFR correlations to constrain feedback. It deserves a serious referee. I would send it to review with a request for error estimates and some test of the half-mass/half-light assumption—if the authors can show the proxy is not differential, the central result stands.","headline":"Useful, honest comparison of the joint mass–SFR–size relation in EAGLE and TNG100; the reported EAGLE-vs-observed slope difference is plausible but rests on an untested half-mass-to-half-light proxy and no error bars.","tokens_in":21648,"tokens_out":3119,"would_cite":true,"duration_ms":28927,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the observed mass–SFR–size relation — the tendency for galaxies to become more compact as they move away from the star-forming main sequence — is reproduced by two independent cosmological simulations, and that its st","keywords":["galaxy stellar mass–size relation","star-forming main sequence","galaxy star formation rate","galaxy size","EAGLE simulation","IllustrisTNG","galaxy evolution","cosmological hydrodynamical simulations"],"falsifier":"Measure the projected half-light radii of the simulated galaxies themselves (or compute the half-mass-to-half-light ratio as a function of ΔlogSFR): if that ratio varies systematically with SFMS offset or redshift, the direct adoption of 3D half-mass radii as a proxy for observed sizes is invalidated and the claimed strength comparison between EAGLE, TNG100, and observations no longer follows. Alternatively, a high-precision observational measurement of the size–offset relation at 1.5 < z < 2.5 that still shows no size decrease away from the SFMS would falsify the claim that the trend persists","tokens_in":20835,"feed_emoji":"🔭","tokens_out":6466,"duration_ms":59854,"temperature":0.7,"pith_summary":"The paper sets out to show that the joint mass–SFR–size relation — the observed tendency for galaxies to become smaller and more compact as they drift away from the ridge of the star-forming main sequence (SFMS) — is not an accidental feature of one dataset but an emergent outcome of cosmological galaxy formation. By comparing SDSS and CANDELS galaxies with the EAGLE and TNG100 simulations across three redshift bins up to z = 2.5, the authors find that both simulations reproduce the trend at all redshifts, whereas the observations only show it clearly below z ≈ 1.5. The two simulations disagree on its strength: EAGLE makes the size drop with SFMS offset too steeply relative to observations, and TNG100 too shallowly. If the paper is right, this relative-size measure becomes a sharp new constraint on how feedback and gas accretion shape galaxies, and the trend's physical origin lies in time-variable star formation — compact galaxies at fixed mass scatter more in SFR over time while having similar average rates.","feed_headline":"Simulations confirm galaxies shrink off the star-forming main sequence","feed_subtitle":"EAGLE makes the size drop too steep, TNG100 too shallow — a new test for galaxy-formation models.","key_machinery":"The analysis is carried by a relative-size metric, ΔlogR_e: each galaxy's logarithmic size is compared with the average size expected at its stellar mass from a power-law mass–size relation fitted to galaxies on the SFMS ridge. This removes the dominant mass dependence and, the authors argue, mostly cancels the systematic differences between simulated 3D stellar half-mass radii and observed projected half-light radii. The partner quantity is the SFMS offset ΔlogSFR, the logarithmic distance of a galaxy above or below the best-fit star-forming main sequence, computed separately for each sample and redshift bin. The paper then compares the median ΔlogR_e versus ΔlogSFR curves for observations,","core_discovery":"In the paper's own terms, the central discovery is that the mass–SFR–size relation is a genuine, persistent feature of galaxy formation physics, not a calibration artifact: both EAGLE and TNG100, which were not tuned to reproduce it, yield the observed decrease of relative galaxy size ΔlogR_e with increasing offset ΔlogSFR from the star-forming main sequence, from z = 0 to z = 2.5. The observed samples show the same 'ridge-largest' pattern at z < 1.5 but not at 1.5 < z < 2.5, which the authors attribute to larger measurement uncertainties at high redshift. Quantitatively, EAGLE predicts a stronger size dependence on SFMS offset than observed at all redshifts, while TNG100 predicts a weaker o","pith_inferences":["If the half-mass-to-half-light ratio of simulated galaxies varies with SFMS offset or redshift, the paper's strength comparison could be biased; a direct check — measuring projected half-light radii of simulated galaxies or computing that ratio as a function of ΔlogSFR — would settle how much of the EAGLE–observation gap is physical rather than definitional.","The compactness–burstiness link suggests a common driver with the fundamental metallicity relation and the scatter of the SFMS: size may serve as a cheap observational proxy for SFR stochasticity in galaxies where direct variability measures are unavailable.","A testable extension is to track individual galaxies in the simulations through the ΔlogR_e–ΔlogSFR plane over time to verify that compact systems actually oscillate between the upper and lower SFMS envelopes; if they instead sit at one offset, the time-variability explanation would need revision.","The z > 1.5 observational null could be tested with rest-optical sizes at high redshift and matched SFR indicators; if the trend still fails to appear, the interpretation would shift from measurement error to a genuine redshift evolution in the coupling."],"forward_implications":["The joint mass–SFR–size relation is an emergent property of cosmological simulations: neither EAGLE nor TNG100 was calibrated to match it, so its presence in both indicates it follows from the standard physics they share.","The observed non-detection at 1.5 < z < 2.5 is likely a measurement-uncertainty effect rather than a physical absence of the trend, since both simulations keep producing it there; higher-precision size measurements at those redshifts should recover it.","The EAGLE–TNG100 difference in trend strength supplies a new observational constraint: any model that wishes to match both the SFMS and the mass–size relation must also reproduce the intermediate slope of the ΔlogR_e–ΔlogSFR relation.","Interpreting the relation through SFR time variability implies that compact and extended galaxies of the same mass and time-averaged SFR can nevertheless be distinguished by their SFR scatter, linking galaxy size to the burstiness of star formation.","The near-flat relation for galaxies above log(M*/M⊙) ≈ 10.8 localizes the mechanism to lower-mass galaxies, where gas cycling and feedback-driven fluctuations are strongest."],"fun_headline_variants":["Simulations confirm galaxies compact off main sequence but strength differs","High-z galaxies show no size drop at last observation, simulations still do","EAGLE and TNG100 both show shrinkage, but with opposite biases","Two simulations confirm compact galaxy trend but magnitude off","Galaxy size drops off main sequence in both simulations, but not equally"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The argument leans on treating the 3D stellar half-mass radius stored in the simulations as a stand-in for the observed projected half-light radius, assuming that any systematic mismatch mostly cancels in the relative-size metric; if the ratio between these radii depends on how far a galaxy sits from the star-forming main sequence or on redshift, the simulated-versus-observed strength of the trend could be biased.","fun_headline_variants_meta":{"raw":{"variants":["Simulations confirm galaxies compact off main sequence but strength differs","High-z galaxies show no size drop at last observation, simulations still do","EAGLE and TNG100 both show shrinkage, but with opposite biases","Two simulations confirm compact galaxy trend but magnitude off","Galaxy size drops off main sequence in both simulations, but not equally"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001468,"raw_usage":{"total_tokens":5766,"prompt_tokens":798,"completion_tokens":4968,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":4879}},"tokens_in":542,"tokens_out":4968,"duration_ms":37718,"temperature":1.0,"reasoning_tokens":4879,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T19:31:14.354301+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the projected half-light radii of the simulated galaxies themselves (or compute the half-mass-to-half-light ratio as a function of ΔlogSFR): if that ratio varies systematically with SFMS offset or redshift, the direct adoption of 3D half-mass radii as a proxy for observed sizes is invalidated and the claimed strength comparison between EAGLE, TNG100, and observations no longer follows. Alternatively, a high-precision observational measurement of the size–offset relation at 1.5 < z < 2.5 that still shows no size decrease away from the SFMS would falsify the claim that the trend persists","supporting_citations":[],"review_version":1}