{"id":"cf93162b-6c73-4b6e-b502-d47f6991009d","arxiv_id":"2510.18293","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Identical subgrid AGN jet feedback produces systematically different lobe morphologies in SPH, moving-mesh, and grid hydro codes, with comparable net ambient heating.","lead":"Astrophysicists implemented the same subgrid AGN jet model in three hydrodynamics codes—SPH, moving-mesh, and grid—and ran idealized jet simulations. The three codes produce systematically different jet lobe shapes, yet have a similar total impact on the surrounding gas.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fixed-neighbour-number injection is a confound: Section 4.4 shows AREPO and PLUTO converge under fixed-neighbour-mass, so the 'intermediate PLUTO' ranking is not a pure solver effect.","rationale":"The reader's weakest assumption is exactly the injection-scheme coupling, and the paper's own Section 4.4 is the strongest evidence that this assumption fails. I agree with the CONDITIONAL verdict, but I want to sharpen the test that would settle the quantitative ordering claim. The paper's central contribution is not merely that 'variations exist'—that is almost expected for different discretizations—but the specific ranking 'SWIFT shorter/wider/hotter; AREPO longer/thinner/cooler; PLUTO intermediate' which is stated in the abstract and reiterated in conclusions. For that ranking to be informative about hydro solvers, the injection scheme must not systematically interact with mass resolution. Section 4.4 demonstrates a strong interaction: switching to fixed neighbour mass removes the AREPO–PLUTO distinction across all three injection velocities (Fig. 9). This means the reported ordering is contingent on the fiducial FNN scheme. The paper does acknowledge this and frames it as part of the coupling dependence, so the qualitative conclusion is internally consistent. However, the quantitative claim (e.g., 'PLUTO lobes are of intermediate length') is not a robust property of the Eulerian method; it is a property of FNN+PLUTO's variable cell masses. The lack of multiple realizations compounds this: random launching angles could shift lobe metrics by an unquantified amount, and the paper reports no error bars. Thus the most load-bearing concern is that, without a coupling-robustness quantification, the headline ordering could mislead readers about intrinsic solver differences. A concrete experiment—fixed-neighbour-mass runs with multiple seeds—would settle whether the ordering persists or collapses.","tokens_in":32119,"tokens_out":4361,"duration_ms":37758,"concrete_test":"Re-run the fiducial uniform-medium simulations (N=10^8, v_j=4×10^4 km/s, t=98 Myr) with the fixed-neighbour-mass injection of §4.4 in all three codes, repeating each with at least 3 random seeds for the launching angles. Compare the distribution of lobe length, width, and temperature at t=98 Myr between AREPO and PLUTO under both injection schemes. If the distributions overlap within seed scatter under fixed-neighbour-mass, the 'PLUTO intermediate' ranking in Fig. 5 is an injection-scheme artifact; if they remain separated, the solver-based ordering survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To isolate the hydrodynamical solver (Abstract, §1), the injection scheme must couple identically across codes. The fiducial scheme injects into the nearest 10 gas elements, mass-weighted (§3.2, Eqs. 6–11). §4.4 shows this is not neutral: PLUTO cells in the jet region reach arbitrarily low masses (§4.2), so the same injected kinetic energy yields higher effective jet velocities than in SWIFT/AREPO. Switching to a fixed neighbour mass (M_ngb = 20×10^5 Msun) makes PLUTO lobes longer/cooler and the AREPO–PLUTO lobe-length tracks converge at all three injection velocities (Fig. 9, middle vs bottom); SWIFT is nearly unchanged. Hence the headline ordering—SWIFT short/wide/hot, AREPO long/thin/cool, PLUTO intermediate—is partly an artifact of the fixed-neighbour-number choice rather than a robust Eulerian-solver property. The paper concedes this (§4.4, §6). Additionally, single realizations and random launching angles provide no error bars, so the magnitude of the claimed code differences is not separated from stochastic scatter.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a virtual-particle jet-launching scheme implemented identically in three hydrodynamical codes—SWIFT (SPH), AREPO (moving-mesh), and PLUTO (static grid)—and uses it to compare idealised, non-relativistic AGN jets and their remnants in uniform and stratified media. The authors report that all codes produce self-similar lobe expansion, with SWIFT lobes shorter/wider/hotter, AREPO lobes longer/thinner/cooler, and PLUTO lobes intermediate, and that all transfer about 60% of injected energy to the ambient medium. They also study resolution, injection-velocity, and neighbour-mass variations, and conclude that code differences arise from the coupling of feedback to resolvable scales and from effective resolution, but are likely subdominant to subgrid-model uncertainty in cosmological simulations.","tokens_in":32401,"tokens_out":4706,"duration_ms":44776,"significance":"If the reported differences are robust, the paper would be a valuable reference for the AGN-feedback community: it uses a deliberately code-agnostic injection model, includes systematic resolution and velocity studies, a stratified-medium setup, and a remnant phase, and it confronts the simulations with an analytic self-similar benchmark that is not fitted to the data. The inclusion of the fixed-neighbour-mass control in §4.4 is a particular strength, as it partially tests the sensitivity of the comparison to the injection coupling. The main value lies not in any single code's prediction but in quantifying the spread across methods and in identifying effective jet velocity and mass resolution as key mediating factors.","major_comments":[{"comment":"The claim in the abstract and §1 that the comparison 'isolates the impact of hydrodynamical solvers' is not supported by the fiducial injection scheme. Virtual particles deposit into the nearest 10 gas elements, but PLUTO cells in the jet region can reach arbitrarily low masses, so the same injected kinetic energy (Eq. 8) produces a higher effective jet velocity in PLUTO than in SWIFT/AREPO. Section 4.4 explicitly shows that with fixed neighbour mass M_ngb = 20×10^5 Msun, AREPO and PLUTO lobe-length tracks converge at all three injection velocities (Fig. 9 bottom vs middle), while SWIFT is nearly unchanged; Section 6 concedes that 'Using a fixed-neighbour-mass scheme reduced differences between these two codes.' Thus the headline ordering SWIFT-short/AREPO-long/PLUTO-intermediate is partly an artefact of the fixed-neighbour-number choice rather than a pure property of the hydrodynamical","section":"§4.4, Fig. 9, Eqs. (6)–(11)"},{"comment":"All lobe metrics are single realisations. Only one random launch-angle sequence is used per run (Section 3.2), and Section 4.4 itself notes sensitivity to 'the stochasticity of individual runs.' Figures 5–7 and 9 contain no error bars or ensemble scatter, so the magnitude of code-to-code differences (e.g., 'PLUTO over-predicts lobe volume by 20%' in §4.1, or the 4–14% resolution trends in §4.2) is not separated from run-to-run variance, which is likely significant given the observed KH and RT instabilities. I ask for at least a small number of repeat runs for the fiducial and for the key neighbour-mass comparison, with reported scatter; if that is impractical, the text should explicitly state that all quantitative differences are single-realization estimates.","section":"§4.1, Figs. 5–7 and 9"},{"comment":"The lobe length and width are defined by ad hoc thresholds (T > 10^8 K or v > 0.5×10^4 km/s) and by averaging the furthest 1% of lobe elements. These thresholds are not varied, and the cross-check with tracer-based definitions is only reported for AREPO and PLUTO, not SWIFT. Since all quantitative claims about code ordering rest on these definitions, the paper should provide a robustness test (e.g., varying the temperature/velocity thresholds and the percentile) or demonstrate explicitly that the code ordering is insensitive to the choice.","section":"§4.1, bullet list 'Jet lobe material'"}],"minor_comments":[{"comment":"The momentum-update formula is taken from Bourne & Sijacki (2017), but the sign convention for p_i0 relative to the jet direction is not defined, and the condition under which 'an additional thermal component' is added is not specified. A short derivation or explanatory sentence would help reproducibility.","section":"§3.2, Eq. (11)"},{"comment":"The parameter v_j is called the 'injection velocity' throughout, but Section 4.1 shows that the effective gas velocity in PLUTO can be higher because of low cell masses. Consider consistently using 'virtual-particle speed' for v_j to avoid confusion.","section":"§3.2 and §4.1"},{"comment":"The point that an equal number of resolution elements 'inherently disadvantages' SWIFT is important, but it appears only parenthetically in the text. It should be stated in the initial-conditions section (Section 3.3) so that the reader can interpret the resolution study from the start.","section":"§4.2"},{"comment":"The two PLUTO resolutions (SR and HR) are both discussed, but the figures and text sometimes refer only to 'red' without always clarifying which run is shown. Explicit labels (e.g., 'PLUTO-SR' and 'PLUTO-HR') in all panels would improve clarity.","section":"§4.5, Fig. 12"},{"comment":"The statement that PLUTO jet simulations 'almost always employ' the special-relativistic or MHD modules is a broad claim without supporting citations. Please add specific references or soften the statement.","section":"§5.1.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about the main injection-coupling confound—it is not a case of concealment—but the abstract's 'isolate the impact of hydrodynamical solvers' overstates what the fiducial setup can show. The fixed-neighbour-mass control and the analytic benchmark are strengths, and the underlying data set is valuable. I would encourage revision rather than rejection: re-frame the central claim as a comparison of coupled solver–injection systems, quantify the coupling effect, and add uncertainty estimates or explicit caveats about single-realization metrics. The lack of repeat runs is the main quantitative weakness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look before you finalize any claims about cross-code jet feedback consistency. The paper implements the same virtual-particle jet injection scheme in SWIFT, AREPO, and PLUTO and compares lobe evolution, energetics, and remnants. The genuinely new piece is that the injection model is code-agnostic and the comparison is systematic: resolution tests, three injection velocities, stratified medium, remnants, and a fixed-neighbour-mass control. The headline result—that identical subgrid prescriptions still produce code-dependent lobes, with SWIFT short/wide/hot, AREPO long/thin/cool, and roughly similar total energy deposition—is supported by the runs and is a useful caution for anyone comparing cosmological simulations that differ in hydro solver.\n\nWhat it does well: the controlled setup, the use of the self-similar analytic solution as an external benchmark rather than a fitted curve, and the unusually candid treatment of the injection scheme's limitations. Section 4.4 explicitly shows that the fiducial fixed-neighbour-number injection is not neutral. PLUTO cells in the jet region reach very low masses, so identical energy injection produces higher effective jet velocities, and switching to a fixed neighbour mass makes AREPO and PLUTO lobe lengths converge across all three injection velocities. The paper concedes this in Section 4.4 and in the conclusions. That honesty is real and deserves credit.\n\nThe soft spots are in proportion. The quantitative ordering—especially 'PLUTO intermediate'—is not a pure solver effect; it depends on the injection coupling. The abstract's summary is a bit cleaner than the evidence, though the conclusions are more careful. Lobe metrics come from single realizations with no error bars, and the lobe definition is threshold-based and ad hoc; the absence of public run artifacts makes independent verification harder. Those are real weaknesses but they affect the magnitude and robustness of the ranking, not the central qualitative conclusion that hydro method alone can change lobe morphology under identical subgrid feedback.\n\nWho is this for? Anyone working on AGN feedback in galaxy formation simulations, particularly those using SWIFT, AREPO, PLUTO, or comparing results across codes. It deserves a serious referee and a request for public data, a few repeat runs to quantify stochastic scatter, and a more balanced presentation of the injection-scheme dependence in the abstract. I would not desk-reject it.","headline":"A careful, honest three-code comparison of a genuinely code-agnostic jet injection scheme; the qualitative morphology ordering holds up, but the paper itself shows that part of that ordering is set by the fixed-neighbour-number injection choice, not by the solver alone.","tokens_in":32945,"tokens_out":1453,"would_cite":true,"duration_ms":16248,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hydro solver sets AGN jet lobe shape; feedback energy stays the same.","keywords":["AGN feedback","jet feedback","subgrid models","SPH","moving-mesh","Eulerian grid","jet lobes","code comparison"],"falsifier":"Run the fixed-neighbour-mass injection in all three codes across the full resolution and jet-velocity grid; if the lobe length, width, and temperature distributions collapse onto each other, the claim that the hydrodynamical method itself drives the differences is falsified. The paper's own lower-resolution runs show better code-to-code agreement, so a resolution-convergence study is the decisive test.","tokens_in":31947,"feed_emoji":"🚀","tokens_out":9533,"duration_ms":74283,"temperature":0.7,"pith_summary":"This paper asks whether the choice of hydrodynamical code changes how AGN jet feedback deposits energy in the gas when the subgrid injection model is exactly identical. The authors build a code-agnostic virtual-particle jet launcher and run it in three codes: an SPH code, a moving-mesh code, and a static-grid code. In uniform and stratified media, all jets drive bow shocks and inflate self-similar lobes, but lobe shape and temperature differ systematically—short, wide, hot lobes in SPH; long, thin, cool lobes in the moving-mesh; intermediate in the grid code. Yet the energy budget is the same: all codes transfer about 60% of the injected energy to the ambient medium, and after switch-off the remnants heat their surroundings almost identically. The paper concludes that the hydrodynamical method, through its effective resolution and how feedback couples to it, is responsible for the morphological differences, and that in cosmological simulations these solver differences are probably smaller than subgrid-model uncertainties.","feed_headline":"Same jet model, three codes: different lobes, same heating","feed_subtitle":"Solver choice reshapes AGN jet lobes, yet all three codes deposit the same energy into the surrounding gas","key_machinery":"The virtual-particle jet injection scheme: pairs of hydrodynamically decoupled particles are launched from the centre in opposite directions within 15-degree cones, travel ballistically to a radius of 10 kpc, and deposit their mass, momentum, and kinetic energy into the 10 nearest gas elements in a mass-weighted fashion. This scheme is deliberately identical across the three codes, making each code's effective resolution the only remaining variable. The self-similar analytic scaling for lobe length, L ~ (P_j / rho0)^(1/5) t^(3/5) in a uniform medium (with a generalisation for power-law atmospheres), serves as the benchmark against which all three codes are judged; SWIFT under-predicts, AREPO","core_discovery":"Central claim: with the subgrid jet prescription held identical, the hydrodynamical discretisation sets lobe morphology. A virtual-particle launcher deposits mass, momentum, and energy into the ten nearest gas elements; in a uniform medium SWIFT produces short, wide, hot lobes, AREPO long, thin, cool lobes, and PLUTO intermediate ones. The same ordering holds in stratified media and in post-switch-off remnants, yet all codes transfer about 60% of the injected energy to the ambient medium and remnant thermodynamic profiles converge. The mechanism is injection coupling: PLUTO's fixed-volume cells reach very low masses in the jet channel, so equal energy injection yields higher effective jet ve","pith_inferences":["A direct consequence the authors leave implicit: reported 'code differences' in AGN feedback studies may often be partly an artefact of injection-region coupling rather than the Riemann solver or discretisation; future comparison projects could standardise both the injected neighbour mass and the time-stepping near the injection region.","If the energetic impact is truly code-independent while lobe morphology is not, then radio-lobe appearance (length, width, temperature) is a poor guide to the feedback energy actually deposited in the gas—a caution for interpreting X-ray cavities and radio observations.","A testable extension: implement the same virtual-particle scheme in a fourth, meshless finite-mass code without mass refinement; the framework predicts its lobe properties should fall between the SPH and moving-mesh results.","The resolution trend—codes agreeing better at low resolution—implies that calibration of subgrid jet models on one code may not transfer to another at high resolution; calibrating per code may be necessary even with identical physics."],"forward_implications":["Identical subgrid jet prescriptions do not guarantee identical jet morphology: solver choice alone shifts lobe length, width, and temperature by tens of percent at fixed resolution and parameters.","The energy budget is code-independent at the level that matters for feedback: all codes deposit roughly 60% of the injected jet energy into the ambient medium, mostly as thermal energy, and leave similar density, temperature, and entropy profiles.","In cosmological simulations, where resolution is coarser and subgrid choices dominate, solver differences in jet morphology are likely subdominant; averaging over galaxy populations would further wash them out.","The way a subgrid model couples to the resolution elements—fixed neighbour number vs fixed neighbour mass—is itself a modelling choice that can change lobe evolution as much as the solver, so feedback schemes should specify and test this coupling."],"fun_headline_variants":["Code choice reshapes jet lobes, not energy output","Same jet recipe, three codes: lobes differ, heating same","Hydro solver decides lobe shape, not energy","Jet lobes vary by code, but heating matches"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's central comparison assumes the virtual-particle injection scheme is truly code-agnostic; the paper itself demonstrates that it is not, because the fixed-neighbour-number deposition produces very different effective jet velocities when a code's resolution elements have very different masses.","fun_headline_variants_meta":{"raw":{"variants":["Code choice reshapes jet lobes, not energy output","Same jet recipe, three codes: lobes differ, heating same","Hydro solver decides lobe shape, not energy","Jet lobes vary by code, but heating matches"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1213,"prompt_tokens":837,"completion_tokens":376,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":314}},"tokens_in":581,"tokens_out":376,"duration_ms":3867,"temperature":1.0,"reasoning_tokens":314,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T08:52:54.378698+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the fixed-neighbour-mass injection in all three codes across the full resolution and jet-velocity grid; if the lobe length, width, and temperature distributions collapse onto each other, the claim that the hydrodynamical method itself drives the differences is falsified. The paper's own lower-resolution runs show better code-to-code agreement, so a resolution-convergence study is the decisive test.","supporting_citations":[],"review_version":1}