{"id":"8058e7a4-6291-4231-8c36-e6d5af9fc7df","arxiv_id":"1908.05019","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Gravitational softening in SPH galaxy simulations must exceed two analytic thresholds to avoid suppressing photo-heating and feedback; below these thresholds, galaxy sizes, stellar masses, and halo concentrations change artificially.","lead":"This paper shows that galaxy formation simulations are much more sensitive to the chosen gravitational softening length than dark-matter-only simulations, and provides simple analytic rules for choosing a safe value. Astronomers running such simulations can use these rules to avoid numerical artifacts that artificially change galaxy sizes, star formation histories, and dark matter halo structure.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equation 13 does not follow from eqs. 10–11: solving nmax_H = nH,tc gives ε_eFB ≈ 1.0 kpc for the Np=376 run, not 0.5 kpc, so the quoted threshold and the 'optimal' softening recommendation are internally inconsistent by about a factor of two.","rationale":"The reader accepted with high confidence and identified generalizability as the weakest assumption. The stress-test found a more immediate internal issue: eq. 13 does not follow algebraically from eqs. 10 and 11 as written. The numerical coefficient in eqs. 12/13 appears to be too small by a factor of about two for both simulation resolutions. This matters because the paper uses ε_eFB as a quantitative criterion to interpret which runs suffer over-cooling and to recommend an 'optimal' softening length. The central qualitative claim—that small softening enhances numerical over-cooling and that hydrodynamical simulations are more softening-sensitive than DM-only runs—is supported by the simulation suite and is not overturned by this arithmetic discrepancy. However, the specific threshold and the practical recommendation in Section 5.2 need to be corrected or explicitly reconciled with eqs. 10 and 11, including any O(1) kernel-volume or cooling-function factors that were absorbed into the order-of-magnitude estimates. A conditional acceptance, pending that reconciliation, is therefore more appropriate than an unconditional accept. No ad hominem is intended; this appears to be a normalization or algebraic slip, but it is load-bearing for the paper's quantitative conclusions.","tokens_in":44456,"tokens_out":39726,"duration_ms":388104,"concrete_test":"Re-derive ε_eFB from eqs. 10 and 11 exactly as written: solve nmax_H = nH,tc for ε with Nngb=58, X=0.75, f=0.1, T=10^7.5 K, and mg=2.26×10^5 M⊙; check whether the result is ≈1.0 kpc rather than ≈0.5 kpc. If it is, recompute the statements in §3.3 and §5.2 that the Np=376 run and the ε_opt=700 pc recommendation lie above ε_eFB. In addition, inspect the original Dalla Vecchia & Schaye (2012) formula for nH,tc to confirm whether its coefficient is normalized to mg=10^5 M⊙ or mg=10^6 M⊙; a sqrt(10) shift in nH,tc would change ε_eFB by an additional factor of about 1.47 and must be stated explicitly.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central feedback-efficiency criterion is derived by demanding nH,tc (eq. 11) ≥ nmax_H (eq. 10). Combining eqs. 10 and 11 with the stated normalizations (Nngb=58, X=0.75, f=0.1, T=10^7.5 K, Plummer-equivalent softening) yields ε_eFB = 350 pc × (180/26)^(1/3) × (mg/10^5 M⊙)^(1/2) ≈ 667 pc × (mg/10^5 M⊙)^(1/2). For the Np=376 run (mg=2.26×10^5 M⊙) this is ≈1.0 kpc, not the 0.5 kpc quoted in §3.3 and eq. 13; for the Np=188 run (mg=1.81×10^6 M⊙) it is ≈2.8 kpc, not 1.5 kpc. No kernel-volume factor or cooling-function normalization is identified in the text that would reconcile this factor of ≈1.9–2.0. The discrepancy is load-bearing: the paper uses ε_eFB to state that EAGLE's high-resolution simulation falls short by only a factor 1.4 and recommends ε_opt = 0.022(L/Np) ≈ 700 pc for the high-resolution runs as 'ensuring feedback will be maximally efficient.' With the corrected coefficient, 700 pc lies below the criterion (≈1.0 kpc), so the recommended softening would fail by the paper's own standard. The qualitative conclusion that small softening enables numerical over-cooling is not at issue; the quantitative threshold and the resulting practical guidance are.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses a suite of cosmological SPH simulations with EAGLE-like subgrid physics, run in a 12.5 cMpc box at two mass resolutions (Np = 188 and 376 per side), to study how the gravitational softening length affects galaxy formation and dark matter halo structure. Softening is varied by factors of two around the fiducial EAGLE values, and the runs include both non-radiative and full-physics cases. The authors derive analytic criteria for choosing the softening: an upper limit from resolving low-mass haloes, a lower limit from two-body collisional heating of gas, a minimum resolved escape velocity (Eq. 8), and a feedback-efficiency threshold (Eqs. 12-13) that prevents numerical over-cooling of supernova-heated gas. They then test these criteria against baryon fractions, cosmic star formation histories, galaxy stellar mass functions, galaxy sizes, and DM circular velocity profiles. The central qualitative conclusion is that hydrodynamical simulations are far more sensitive to softening than dark-matter-only simulations, with too-small softening suppressing photo-heating, promoting over-cooling, and exacerbating mass segregation, sometimes causing galaxy sizes to increase as softening decreases.","tokens_in":44859,"tokens_out":15152,"duration_ms":153873,"significance":"If the quantitative criteria are correct, this is an important contribution: it provides simple, physically motivated rules of thumb for a community that routinely adopts softening lengths without such tests, and it demonstrates the non-trivial coupling between numerical force resolution and subgrid feedback. The study is strengthened by the systematic factor-of-two coverage of softening, by the use of two mass resolutions, and by direct diagnostics such as baryon fractions and stellar birth densities that make the claimed effects falsifiable. The qualitative results, especially the v_epsilon criterion and the identification of mass segregation as a softening-dependent effect, are valuable and likely robust. However, the central feedback-efficiency threshold contains an arithmetic error of a factor of about 1.9, which directly affects the paper's quantitative statements and its recommended 'optimal' softening. The paper therefore requires a major revision before the practical guidance can be accepted.","major_comments":[{"comment":"The coefficient in Eq. (12) does not follow from equating Eqs. (10) and (11). With the stated normalizations, n_H,tc = n_max^H implies epsilon_eFB = 350 pc x (180/26)^(1/3) x (Nngb/58)^(1/3) x (X/0.75)^(1/3) x (f/0.1)^(-1) x (T/10^7.5 K)^(-1/2) x (mg/10^5 Msun)^(1/2), approximately 667 pc x (mg/10^5 Msun)^(1/2). The missing factor (180/26)^(1/3) is about 1.9 and is not a convention effect: Eq. (10) already refers to the Plummer-equivalent softening, with the eps_sp = 2.8 eps conversion folded into the 180 cm^-3 calibration. Consequently, the values quoted in Section 3.3 (epsilon_eFB about 0.5 kpc and 1.5 kpc for the Np = 376 and Np = 188 runs) should be about 1.0 kpc and 2.8 kpc. This is not a cosmetic issue: Sections 5.1(iii) and 5.2 use these values to state that EAGLE's high-resolution run falls short of epsilon_eFB by only a factor of 1.4 and that the recommended softening 'ensure[s] that feedback will be maximally efficient.' Under the corrected criterion, 700 pc is below epsilon_eFB, and the shortfall is about 2.9, not 1.4. Please recompute the threshold, revise the numerical statements throughout, and reassess the practical recommendations.","section":"Section 3.3, Eqs. (10)-(13)"},{"comment":"The recommendation epsilon_opt about 0.022 L/Np (approximately 700 pc for Np = 376 and 1400 pc for Np = 188) is presented as ensuring efficient feedback, but with the corrected epsilon_eFB from the previous comment it falls below the paper's own feedback-efficiency threshold for both resolutions. The statement in Section 5.2 that 'these criteria ensure that feedback will be maximally efficient' is therefore not supported by the corrected algebra. Because epsilon_eFB and the DM-only convergence radius rconv about 0.055 L/Np (about 1.8 kpc for Np = 376) are both constraints, the paper should discuss the resulting allowed window (roughly 1.0-1.8 kpc for the high-resolution runs) and explain whether any of the simulated configurations actually satisfies both requirements, rather than simply doubling the fiducial softening.","section":"Section 5.2, optimal softening recommendation"}],"minor_comments":[{"comment":"The text quotes 'about 27 per cent for epsilon = 700 pc' for stars formed at z above about 11.5 and then, in the same paragraph, quotes 'only about 9 per cent of all stars' for the same run; please state explicitly that these refer to different stellar populations (pre-reionization versus all epochs) to avoid an apparent contradiction.","section":"Section 4.1.3, first paragraph"},{"comment":"The phrase 'the upper middle-right panel' is ambiguous; please specify the panel corresponding to epsilon0 = 700 pc.","section":"Figure 7 caption"},{"comment":"There are minor language errors: 'eﬀected' should be 'affected' in Section 2.2, and 'we also with to acknowledge' should be 'we also wish to acknowledge' in the Acknowledgements.","section":"Section 2.2 and Acknowledgements"},{"comment":"The derivation of n_H,tc in Eq. (11) is quoted from Dalla Vecchia & Schaye (2012), but the text would benefit from a brief statement of the adopted cooling function and the kernel-volume normalization, since the coefficient 26 cm^-3 enters directly into the corrected feedback threshold.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The qualitative findings are valuable and the simulation suite is well suited to the question, but the factor-of-two error in Eq. (12) affects the central quantitative threshold and the recommended softening in Section 5.2. I would ask the authors to correct the algebra and re-evaluate the affected statements; with that revision, the paper could become acceptable. I would not recommend acceptance in the current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The central empirical claim is convincing: hydrodynamical runs have a strong, physically intelligible sensitivity to gravitational softening, unlike DM-only runs. The paper earns its keep with a clean suite of EAGLE-like runs at two resolutions, softening varied by factors of two, and with analytic criteria that are tested rather than just asserted. The v_epsilon escape-speed criterion is verified by the reionization-suppression comparison (Figure 2), and the mass-segregation/size story — smaller softening makes galaxies bigger, not smaller — is a genuinely counterintuitive result worth having.\n\nThe soft spot is real, though. The feedback-efficiency threshold ε_eFB is derived by demanding n_H,tc ≥ n_H,max from eqs. 10 and 11. Doing that algebra gives ε_eFB ≈ 667 pc × (mg/10^5 M_sun)^1/2 with the paper's own normalizations, not the 350 pc quoted in eq. 12. For the Np=376 run (mg=2.26×10^5) that is ≈1.0 kpc, not 0.5 kpc; for Np=188 it is ≈2.8 kpc, not 1.5 kpc. I can't find a kernel-volume or cooling-normalization factor in the text that closes the factor of ≈1.9. This is load-bearing: the paper uses ε_eFB to say EAGLE's high-resolution run falls short by only 1.4 and recommends ε_opt≈0.022(L/Np)≈700 pc as 'ensuring feedback will be maximally efficient.' With the corrected coefficient, 700 pc is below the threshold, so by the paper's own standard the recommended setup would not do the job.\n\nThat said, the qualitative conclusion — small softening opens the door to numerical over-cooling — is not in doubt; Figure 3 shows the same trend. The limitations are mostly acknowledged: one realization per parameter point, EAGLE-specific subgrid physics, no public code. The mass-segregation interpretation is called speculative by the authors themselves. So the balance is good; this is a genuine contribution with one significant arithmetic slip in the central analytic result.\n\nI'd send it to a serious referee. The simulation campaign and the empirical findings deserve publication, but I'd require the derivation of eqs. 12–13 to be corrected and the softening recommendation re-evaluated before acceptance. If that is fixed, this is a paper I'd want in the literature.","headline":"A strong, useful convergence study whose headline feedback criterion (eq. 13) is off by a factor of about two — the qualitative message survives, the quantitative recommendation does not.","tokens_in":45378,"tokens_out":6847,"would_cite":true,"duration_ms":65798,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hydrodynamical galaxy simulations are far more sensitive to the gravitational softening length than dark-matter-only runs are, and this paper derives analytic thresholds that tell when a chosen softening will distort the results.","keywords":["gravitational softening","numerical convergence","SPH simulations","galaxy formation","dark matter haloes","reionization","stellar feedback","galaxy sizes"],"falsifier":"Keep particle masses and subgrid physics fixed and rerun one cosmological volume with present-day softening below $19\\,{\\rm pc}$ and above $156\\,{\\rm pc}$, the paper's $v_\\epsilon$ bounds for its high-resolution case: if the median baryon fraction of $\\sim10^8\\,\\mathrm{M}_\\odot$ haloes at $z=10$ does not jump from near the cosmic mean to a small fraction of it, the minimum-resolved-escape-speed mechanism fails. The same experiment should also show a high-redshift star-formation peak and late-time suppression for the small softening; their absence would falsify the over-cooling story.","tokens_in":44301,"feed_emoji":"🔭","tokens_out":9626,"duration_ms":92534,"temperature":0.7,"pith_summary":"This paper establishes that the gravitational softening length is a first-order numerical parameter in hydrodynamical galaxy formation simulations, not a small detail. In a suite of cosmological smoothed-particle-hydrodynamics runs that keep particle masses and subgrid physics fixed and vary only softening, the star formation history, the abundance of low-mass galaxies, galaxy sizes, and the inner structure of dark-matter haloes all shift systematically as softening is changed by factors of two. The authors derive two analytic thresholds that mark where the damage begins: a minimum resolved escape speed set by gas self-binding, and a critical softening below which thermal feedback radiatively overcools. These thresholds matter because they give simulators concrete numbers to choose softening by, and they warn that a higher-resolution run is not automatically a more trustworthy one.","feed_headline":"Smaller softening does not mean a better galaxy simulation","feed_subtitle":"A convergence study derives the softening thresholds that keep star formation, galaxy sizes, and dark-matter haloes trustworthy.","key_machinery":"The central machinery is a pair of analytic thresholds written in terms of the Plummer-equivalent softening $\\epsilon$ and gas particle mass $m_g$: the minimum resolved escape speed $v_\\epsilon = \\sqrt{2G m_g/\\epsilon}$ (Eq. 8) and the feedback-efficiency softening $\\epsilon_{\\rm eFB}$ that keeps the maximum resolved gas density, $n_{\\rm H}^{\\rm max}\\propto \\epsilon^{-3}$, below the critical density $n_{\\rm H,tc}$ required for thermal feedback to act before cooling (Eq. 13). These thresholds connect force resolution to gas physics, and the paper verifies them against two distinct diagnostics: the baryon fractions of low-mass haloes at $z=10$ and the distribution of gas densities at the moment stars are born.","core_discovery":"On its own terms, the paper claims that hydrodynamical simulations of galaxy formation are far more sensitive to gravitational softening than dark-matter-only simulations, and that the sensitivity is predictable. Softening sets a minimum resolved escape speed $v_\\epsilon = \\sqrt{2 G m_g / \\epsilon}$ for gas particles; once $v_\\epsilon$ exceeds roughly $10\\,{\\rm km\\,s^{-1}}$, the photo-heating associated with reionization is suppressed, low-mass haloes retain baryons, and stars form in systems that should have been quenched. Softening also caps the maximum resolved gas density roughly as $n_{\\rm H}^{\\rm max}\\propto \\epsilon^{-3}$; when that cap pushes gas past the critical density $n_{\\rm H,tc}$ for efficient thermal feedback, a large share of star-forming gas radiates away feedback energy before it can act. In addition, because dark-matter particles are more massive than star particles, small softening accelerates two-body mass segregation, which inflates galaxy sizes and contracts dark-matter haloes even after the baryons responsible for the contraction have been scattered away. The paper's recommended softening---about $0.05$ times the mean inter-particle spacing in comoving units at early times and $0.022$ times that spacing in physical units at late times---keeps circular velocity profiles converged to roughly 10 per cent outside the dark-matter convergence radius.","pith_inferences":["The same two thresholds should apply, at least approximately, to any subgrid feedback model that heats gas to roughly $10^{7.5}\\,\\mathrm{K}$, because the over-cooling criterion depends only on particle mass, softening, and post-heating temperature.","A testable extension of the mass-segregation result: runs with equal-mass dark-matter and star particles should show much weaker softening-driven size inflation, since the paper ties the effect to the mass ratio $\\mu = m_{\\rm DM}/m_{\\rm gas}$.","The softening-dependent star-formation peak near $z_{\\rm reion}$ implies that high-redshift star formation rates from simulations with very small softening deserve suspicion even when their $z=0$ galaxy populations look reasonable.","If these criteria generalize to moving-mesh or adaptive-mesh codes, simulators would gain a common vocabulary for choosing force resolution across methods; the paper does not test this, but the criteria are phrased in resolution-element terms that carry over."],"forward_implications":["Runs with softening below the $v_\\epsilon$ threshold overproduce stars in small haloes after reionization, so their low-mass galaxy stellar mass functions become too steep.","Runs with softening below $\\epsilon_{\\rm eFB}$ convert a large part of star formation into numerically over-cooled events with weak feedback, producing galaxies that are initially too compact.","Runs with softening too large suppress dwarf-galaxy formation because the maximum resolved gas density falls below the star-formation threshold.","At the recommended optimal softening, cosmic star formation histories and stellar mass functions converge across two mass resolutions, and dark-matter circular velocity profiles agree to within about 10 per cent outside the dark-matter convergence radius.","Small softening also alters dark-matter structure: haloes that are baryon-free at $z=0$ can still contain steep central cusps left by baryons that were later scattered away, so dark-matter-only convergence criteria do not carry over to hydrodynamical runs."],"supporting_citations":[{"why":"Supplies the calibrated subgrid model and the fiducial softening lengths that all runs vary around.","marker":"Schaye et al 2015"},{"why":"Defines the stochastic thermal feedback scheme and the critical density n_H,tc below which feedback is numerically efficient.","marker":"Dalla Vecchia & Schaye (2012)"},{"why":"Provides the radiative cooling and heating rates that the full-physics runs use and that make over-cooling a concern.","marker":"Wiersma et al. 2009"},{"why":"Documents how feedback-calibrated simulations produce realistic stellar mass functions and size-mass relations, and warns that high-density star formation yields overly compact galaxies.","marker":"Crain et al. 2015"},{"why":"Provides the dark-matter-only convergence radius r_conv and the mass-segregation/energy-equipartition analysis for unequal-mass particles that the size results build on.","marker":"Ludlow et al. (2019)"},{"why":"Establishes the classical dark-matter-only convergence criteria that the paper contrasts with the hydrodynamical behaviour.","marker":"Power et al. 2003"}],"fun_headline_variants":["Smaller softening inflates galaxies and shrinks haloes","Why smaller softening can ruin galaxy simulations","Softening below a threshold silences galactic feedback","Convergence study finds softening sweet spot for galaxies","Galaxy sizes grow as gravitational softening shrinks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative thresholds are derived and tested for smoothed-particle-hydrodynamics runs with a single calibrated subgrid model, equal numbers of dark-matter and gas particles, and identical softening for all particle species; other numerical setups may shift the numbers or weaken the effects.","fun_headline_variants_meta":{"raw":{"variants":["Smaller softening inflates galaxies and shrinks haloes","Why smaller softening can ruin galaxy simulations","Softening below a threshold silences galactic feedback","Convergence study finds softening sweet spot for galaxies","Galaxy sizes grow as gravitational softening shrinks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000304,"raw_usage":{"total_tokens":1850,"prompt_tokens":1154,"completion_tokens":696,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":770,"completion_tokens_details":{"reasoning_tokens":624}},"tokens_in":770,"tokens_out":696,"duration_ms":7520,"temperature":1.0,"reasoning_tokens":624,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:25:22.017866+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Keep particle masses and subgrid physics fixed and rerun one cosmological volume with present-day softening below $19\\,{\\rm pc}$ and above $156\\,{\\rm pc}$, the paper's $v_\\epsilon$ bounds for its high-resolution case: if the median baryon fraction of $\\sim10^8\\,\\mathrm{M}_\\odot$ haloes at $z=10$ does not jump from near the cosmic mean to a small fraction of it, the minimum-resolved-escape-speed mechanism fails. The same experiment should also show a high-redshift star-formation peak and late-time suppression for the small softening; their absence would falsify the over-cooling story.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the calibrated subgrid model and the fiducial softening lengths that all runs vary around."},{"cited_title":"F., Jenkins A., Frenk C","cited_arxiv_id":null,"evidence_quote":"Establishes the classical dark-matter-only convergence criteria that the paper contrasts with the hydrodynamical behaviour."}],"review_version":1}