{"id":"3cd95515-5a8d-4157-b840-4bf5cec2968a","arxiv_id":"2505.04509","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Both NewHorizon and Illustris TNG50 fail to reproduce the structural diversity of observed low-mass galaxies, falling at opposite extremes of the observed trends.","lead":"Dwarf galaxies made by two leading cosmological simulations look very different from each other and from real dwarfs in a deep sky survey. One simulation makes them too diffuse and clumpy, while the other makes them too compact and concentrated.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Completeness test validates detection only, not the COSMOS2020 photo-z/mass selection; the 'full diversity' claim may rest on an unverified selection function.","rationale":"The paper's main comparison - that NewHorizon and TNG50 bracket the observed COSMOS dwarfs and both deviate from them - is a well-controlled and valuable result, supported by consistent trends across Sersic and non-parametric metrics and by rest-frame checks. However, the headline claim that the simulations 'fail to capture the full diversity' depends critically on the observed sample being a complete census of the dwarf population. The reader's weakest-assumption analysis correctly identified this vulnerability. I agree with the reader's positive assessment, but I would sharpen the concern: the completeness test in Section 3.3 only validates detection by the segmentation algorithm, not the full COSMOS2020 photometric-redshift and stellar-mass selection that defines the observed sample. This is a more specific gap than 'injecting mocks from the simulations under test' alone, because even a detected galaxy can be lost to photo-z or mass quality cuts. The circularity the reader noted is real: the injected galaxies define what 'typical' dwarfs look like, so a population that is absent from both simulations cannot be probed. The proposed test - running injections through the complete selection pipeline and adding still more diffuse mocks - would settle whether the concern lands. If such a test showed that diffuse dwarfs are recovered at high efficiency, the central claim would stand. Until then, the 'full diversity' statement is conditional on an unvalidated selection function, so I recommend CONDITIONAL acceptance rather than unconditional ACCEPT.","tokens_in":29897,"tokens_out":7660,"duration_ms":77879,"concrete_test":"Inject a suite of mock dwarf galaxies - including the NewHorizon and TNG50 mocks used in the paper plus artificially constructed, even more diffuse and fainter Sersic profiles (n < 1, large Reff) - into the HSC-SSP COSMOS deepCoadd images, then run the full COSMOS2020-style photometric-redshift and stellar-mass pipeline (e.g., LEPHARE on the same UV-to-IR photometry) and apply selection criteria (i)-(iii). Measure the recovery fraction after all cuts as a function of intrinsic Reff, Sersic index, and surface brightness. If the recovery fraction declines sharply for diffuse, low-mass objects, the observed sample is incomplete and the 'full diversity' claim must be weakened or corrected. As a complementary check, compare the COSMOS sample against deeper space-based imaging in the same field (e.g., HST/CANDELS or JWST) to test whether a missing diffuse population is present.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that both simulations 'fail to capture the full diversity of COSMOS dwarfs' depends on the COSMOS sample being a fair census of dwarfs in the stated mass and redshift range. The completeness analysis in Section 3.3 and Figure 2 tests only whether injected simulated galaxies are detected by the photutils detection and segmentation procedure on HSC deepCoadd images. It does not propagate those detections through the COSMOS2020 photometric-redshift and stellar-mass selection actually used to build the observed sample: 0.05 < z < 0.25 with fractional redshift error below 10 per cent, and 10^7.5 < M_star/M_sun < 10^9.5 (criteria i-iii of Section 3.3). Faint, diffuse dwarfs - precisely the population that might discriminate between the simulations - are likely to have noisier photometry, larger photo-z uncertainties, and less secure mass estimates, so they could be preferentially rejected by these cuts even when they are detected. Because the injected mocks are drawn from the very simulations whose realism is under test, any population absent from both simulations is invisible to the completeness calibration. The statement in Section 3.3 that completeness is 'expected to be high' therefore overreaches: it validates detection completeness for simulated galaxies, not sample completeness for real dwarfs. If the true dwarf population includes a substantial population that is more diffuse or fainter than NewHorizon, or whose photo-z are unreliable, the observed 'diversity' would be underestimated, and the conclusion that simulations fail to capture it could be an artifact of selection. This directly undermines the headline claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Using ultra-deep HSC-SSP imaging of the COSMOS field, the authors compare Sérsic and non-parametric structural measurements of 1320 dwarf galaxies (10^7.5 < M*/Msun < 10^9.5, 0.05 < z < 0.25) with redshift- and mass-matched mock observations of dwarfs from the TNG50 and NewHorizon simulations. Synthetic images are produced with SED evolution, dust attenuation, HSC PSF convolution, and injection into real HSC backgrounds, with detection performed consistently for observed and simulated galaxies. The central finding is that NewHorizon and TNG50 lie at opposite extremes of the observed structural trends and that both simulations fail to capture the full diversity of the COSMOS dwarfs at lower masses, with better agreement near 10^9.5 Msun. The paper attributes the differences to distinct ISM and feedback implementations.","tokens_in":30156,"tokens_out":4572,"duration_ms":42873,"significance":"If the conclusions hold, the paper provides a stringent, parameter-free test of two state-of-the-art simulation codes in a mass regime where galaxy formation models remain poorly constrained, and it demonstrates a repeatable forward-modeling pipeline for low-mass galaxy morphology. The strength of the paper is its careful matching of samples, consistent PSF treatment, source injection, and the use of rest-frame checks to separate observational bias from intrinsic differences. The comparison to observations is not fitted to the simulations, so the reported discrepancies are informative for feedback physics. The main residual uncertainties are the completeness of the observed sample and the lack of formal statistical tests, both of which are addressable.","major_comments":[{"comment":"The completeness analysis tests only whether injected simulated galaxies are recovered by the photutils detection/segmentation on HSC deepCoadd images; it does not propagate the galaxies through the COSMOS2020 LePhare photometric-redshift and stellar-mass selection (criteria i-iii). The statement that completeness is 'expected to be high' and the inference that the COSMOS sample is an unbiased census of dwarfs in the stated mass and redshift range therefore overreach. If real dwarfs are fainter or more diffuse than both simulations, or have noisier photometry yielding larger photo-z errors, they could be preferentially rejected by the redshift and mass cuts even when detected. I request either an injection run that includes the full photo-z/mass selection or a softened statement of the 'full diversity' claim.","section":"Section 3.3, Figure 2"},{"comment":"No formal two-sample significance tests are reported anywhere in Section 4. The narrative repeatedly uses 'significant' (e.g., 'significantly larger sizes' in Section 4.1; 'highly significant differences' in Section 5.1.2) without a statistical measure. Because the central claim is that the simulations do not reproduce the observed distribution, the paper should quantify the agreement or disagreement using a test such as a two-dimensional Kolmogorov-Smirnov or energy-distance statistic applied to the mass-matched samples, with bootstrap confidence intervals on the medians and distribution widths.","section":"Section 4"},{"comment":"The mock galaxies are drawn from a single snapshot at z approximately 0.2 and then assigned redshifts matching the observed distribution. The authors argue that structural evolution between z=0.25 and 0.05 is small compared with the simulation differences, but no quantitative justification (e.g., a comparison of two snapshots) is given. Since the observed sample spans this full range and the non-parametric metrics are redshift-sensitive, as shown by the rest-frame appendix, a check of structural stability across the snapshot would strengthen the comparison.","section":"Section 3.1.2"}],"minor_comments":[{"comment":"The variable 'Gsini' should be 'Gini'.","section":"Section 4.2"},{"comment":"The word 'redshft' should be 'redshift'.","section":"Figure 3 caption"},{"comment":"The passage beginning 'A 256x256 pixel bin third-order sky correction task...' ends with an incomplete sentence: 'objects smaller than this scale...' This fragment should be integrated into the previous sentence.","section":"Appendix A"},{"comment":"The violin plots would benefit from labeling sample sizes and from adding units to the Reff axis, as is done elsewhere in the paper.","section":"Figure 10"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-suited for MNRAS and the forward-modeling pipeline is a genuine strength. The main risk is that the completeness claim overreaches; if the authors either run the injected mocks through the actual photo-z/mass selection or soften the 'full diversity' conclusion, the central result would be on much firmer ground. The addition of formal statistical tests would also address the main quantitative weakness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read on this one.\n\nThe paper is worth your time. It forward-models TNG50 and NewHorizon dwarf galaxies through a realistic HSC-like pipeline—PSF convolution, injection into real COSMOS backgrounds, matched mass and redshift samples—and compares Sersic and non-parametric morphology against ultra-deep HSC-SSP imaging. The headline result, that the two simulations bracket the observed dwarfs at low masses (NewHorizon too diffuse, TNG50 too compact) and both miss the observed diversity, is cleanly presented. That is a useful, non-obvious target for feedback and ISM physics, and the pipeline is a good template for Rubin/Euclid-era work. The paper is honest about many limitations, including PSF effects and the choices in generating mock images.\n\nThe main caveat, which the stress-test note correctly identifies, is the completeness analysis. The Section 3.3 detection test validates whether injected simulated galaxies are recovered by the photutils detection/segmentation on HSC deepCoadds. But the observed sample is built from COSMOS2020 photometric redshifts and stellar masses, with cuts on redshift range and fractional redshift error. The completeness test does not propagate the synthetic sources through those photo-z and mass selections. Faint, diffuse dwarfs—precisely the population that could decide between the simulations—may be detected yet have noisy photo-z or loose mass estimates and get rejected. And because the injected galaxies are drawn from the two simulations under test, any population absent from both is invisible to the calibration. This is a genuine soft spot. I don't think it kills the paper. The central bracketing result is large in amplitude and likely robust. But the claim that both simulations 'fail to capture the full diversity' is less secure if the observed sample itself is missing the most diffuse dwarfs. The paper says completeness is 'expected to be high' and argues the observed Sersic parameters lie between the simulations; that is suggestive but not a full selection-function test.\n\nOther smaller quibbles: no formal significance tests (K-S or similar) on the distribution comparisons; the single-snapshot approximation for the simulations is justified in one sentence and could be off at the low-mass end; and the physical-driver discussion is speculative, though the authors label it as such.\n\nBottom line: this deserves a serious referee. The useful core—a careful, reproducible forward-modeling comparison that gives the community a clear discrepancy to explain—is solid, and the caveats are addressable in revision. I'd probably take it to the reading group too.","headline":"Solid forward-modeling comparison of dwarf structure in TNG50 vs NewHorizon that makes a useful point about simulation physics, though the sample completeness analysis does not cover the photo-z/mass selection and deserves scrutiny.","tokens_in":30909,"tokens_out":3515,"would_cite":true,"duration_ms":33529,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that NEWHORIZON and TNG50 produce dwarf galaxies with structural properties at opposite extremes of observed COSMOS dwarfs, and that neither simulation captures the full diversity of low-mass dwarfs.","keywords":["dwarf galaxies","galaxy structure","Sersic profiles","Gini-M20 morphology","cosmological simulations","synthetic observations","galaxy feedback","HSC-SSP"],"falsifier":"A direct test would be a survey in the same redshift window reaching roughly 32 mag arcsec$^{-2}$ in the $i$ band, with a completeness function measured by injecting real ultra-diffuse galaxies rather than simulation galaxies. Recovered dwarfs at those depths whose structural distribution remains between the two simulated extremes would support the paper; a recovered distribution matching either simulation's extreme would show the COSMOS baseline was incomplete.","tokens_in":29727,"feed_emoji":"🌌","tokens_out":10668,"duration_ms":107517,"temperature":0.7,"pith_summary":"At the heart of this paper is a simple question: can current cosmological simulations reproduce the diversity of dwarf galaxy structures seen in deep imaging? The authors render mock HSC-like images of dwarfs from the NEWHORIZON and TNG50 simulations, inject them into real backgrounds from the HSC-SSP COSMOS field, and measure the same structural statistics used on 1,320 observed dwarfs in the mass range $10^{7.5}<M_\\star/M_\\odot<10^{9.5}$. They find that the two simulations bracket the observed population from opposite sides: NEWHORIZON dwarfs are too diffuse, extended, and shallow-profiled, while TNG50 dwarfs are too compact, concentrated, and steep-profiled. Only toward $M_\\star\\sim10^{9.5}\\,M_\\odot$ does agreement become comparable, and even there TNG50 remains overly compact. A sympathetic reader should take away that dwarf structure is a discriminating test of star-formation and feedback physics, and that neither model currently passes it at the low-mass end.","feed_headline":"Simulated dwarf galaxies miss the real diversity of shapes","feed_subtitle":"One simulation makes dwarfs too diffuse, the other too compact; the gap closes only near ten billion solar masses.","key_machinery":"The machinery is a matched synthetic-observation pipeline. Simulated galaxies are built from star particles via stellar-population SEDs, dust attenuation, and PSF convolution, then injected into HSC-SSP COSMOS backgrounds so they experience the same detection, segmentation, sky, and measurement biases as observed galaxies; observed and mock samples are matched in stellar mass and redshift. Structural comparisons then use two families of measures: single-component Sérsic fits (effective radius, surface brightness at the effective radius, Sérsic index, ellipticity) and non-parametric statistics (Gini, $M_{20}$, concentration, asymmetry). A rest-frame control at fixed physical scale separates intrinsic simulation properties from observational smearing. This design is what lets the authors attribute residual disagreements to physics rather than to selection effects or the telescope.","core_discovery":"The central discovery claimed is negative in form: neither of the two simulations reproduces the structural range of observed low-mass dwarfs, and the failures are systematic and opposite. In Sérsic terms, NEWHORIZON dwarfs have large effective radii and low Sérsic indices, while TNG50 dwarfs have small effective radii and high Sérsic indices; non-parametric Gini, $M_{20}$, asymmetry, and concentration measurements place NEWHORIZON as clumpy and asymmetric and TNG50 as smooth and overconcentrated. The observed COSMOS dwarfs sit between these extremes, with relatively flat trends of structure with stellar mass, whereas both simulations show stronger mass dependence. The authors rule out their measurement pipeline and the HSC PSF as the cause: rest-frame measurements at fixed physical scale make TNG50's compactness more extreme once PSF smearing is removed, and detection-injection tests show high completeness for the observed sample. They interpret the split as a fingerprint of the sub-grid physics, with NEWHORIZON's bursty, locally coupled supernova feedback evacuating central gas and TNG50's smoother ISM and feedback model concentrating star formation in the center.","pith_inferences":["If the bracketing pattern generalizes, a third simulation with intermediate sub-grid choices should produce dwarfs whose structural distribution falls inside the observed COSMOS cloud; locating that 'Goldilocks' model is a direct target for future simulation comparisons.","The completeness test, which injects the two simulations' own galaxies, cannot detect a population of real dwarfs that is fainter or more diffuse than both; if such a population exists, the observed COSMOS distribution would be incomplete and the true diversity gap would be even larger than reported.","The same matched-injection methodology applied to environment-ranked subsamples could separate feedback-driven from environment-driven structural scatter; the paper's own environmental argument suggests this is testable.","If star-formation burstiness is the culprit, the scatter in structural parameters within each simulation should correlate with the burstiness of individual dwarfs' star-formation histories, a testable prediction the paper gestures toward for its companion analysis."],"forward_implications":["Below $M_\\star\\sim10^{9.5}\\,M_\\odot$, neither simulation's raw structural distributions should be treated as predictions of dwarf morphology; the matched-injection transform is required before comparison.","The direction of the mismatch is tied to ISM and supernova feedback prescriptions, so dwarf structure can discriminate between such recipes.","Rest-frame results imply that TNG50 dwarfs are intrinsically too compact and not merely PSF-biased.","Better agreement at the high-mass end means the discrepancy is specific to the low-mass dwarf regime, where feedback physics is most sensitive."],"supporting_citations":[{"why":"Supplies the NEWHORIZON simulation, its sub-grid physics, and one of the two mock dwarf populations.","marker":"Dubois et al. 2021"},{"why":"Provides the TNG50 simulation run whose dwarf population is the other mock population.","marker":"Nelson et al. 2019"},{"why":"Provides TNG50's galaxy formation model and the feedback and ISM choices under test.","marker":"Pillepich et al. 2019"},{"why":"COSMOS2020 catalogue supplies photometric redshifts, stellar masses, and the observed dwarf sample.","marker":"Weaver et al. 2022"},{"why":"HSC-SSP DR3 deepCoadd imaging provides the observed backgrounds and postage stamps.","marker":"Aihara et al. 2022"},{"why":"The synthetic-image generation procedure is adapted from this method.","marker":"Martin et al. 2022"},{"why":"Provides the software and definitions used to measure Sérsic, Gini-M20, and CAS structural parameters.","marker":"Rodriguez-Gomez et al. 2019"},{"why":"Defines the Gini and M20 statistics used to compare light distributions.","marker":"Lotz et al. 2004"},{"why":"Defines the concentration and asymmetry statistics used in the comparison.","marker":"Conselice et al. 2003"}],"fun_headline_variants":["Simulated dwarfs are too diffuse or too compact","Neither simulation captures real dwarf galaxy range","Simulations fail to reproduce low-mass galaxy structures","Dwarf shapes: simulation extremes miss observed middle"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on treating the COSMOS sample as a fair view of the true dwarf population, but the completeness correction uses mock galaxies from the very simulations whose realism is on trial, so a real population fainter or more diffuse than either simulation could be missing and the conclusion would shift.","fun_headline_variants_meta":{"raw":{"variants":["Simulated dwarfs are too diffuse or too compact","Neither simulation captures real dwarf galaxy range","Simulations fail to reproduce low-mass galaxy structures","Dwarf shapes: simulation extremes miss observed middle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00093,"raw_usage":{"total_tokens":4061,"prompt_tokens":1104,"completion_tokens":2957,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":2899}},"tokens_in":720,"tokens_out":2957,"duration_ms":20195,"temperature":1.0,"reasoning_tokens":2899,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:26:25.390258+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be a survey in the same redshift window reaching roughly 32 mag arcsec$^{-2}$ in the $i$ band, with a completeness function measured by injecting real ultra-diffuse galaxies rather than simulation galaxies. Recovered dwarfs at those depths whose structural distribution remains between the two simulated extremes would support the paper; a recovered distribution matching either simulation's extreme would show the COSMOS baseline was incomplete.","supporting_citations":[],"review_version":1}