{"id":"c779d1a1-06b6-4db2-80a2-1ea6e4acaa86","arxiv_id":"1908.05276","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Relaxed-FOF halo finding plus per-step potential gradient descent makes low-resolution FastPM simulations match high-resolution N-body halo and matter statistics closely enough for survey mocks.","lead":"A fast cosmological simulation technique is tweaked so that small dark matter halos and small-scale matter clustering match slower, high-resolution N-body simulations. The new tools, a mass-dependent halo finder and an in-simulation correction step, could make survey-scale mock catalogs for DESI and LSST affordable.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No out-of-sample test: relaxed-FOF and PGD parameters are fit to TNG300-2-Dark and validated only on that same initial density field, so transfer to DESI/LSST mocks is the untested load-bearing step.","rationale":"The reader identified the same lack of out-of-sample validation; I agree. This is load-bearing because the paper's own validation metrics (missed halo fraction, cross-correlation coefficient, power-spectrum ratios) are all measured on the same simulation pair used for calibration. The matched initial conditions make the halo-by-halo comparison possible, but they also mean the quoted scatter bands are not a test of how the fitted functions behave for a different realization. I do not see an internal mathematical contradiction in the method; the equations are plausible and the paper compares several independent statistics, including RSD and halo-matter cross-power, which are not directly fitted. Those give useful evidence that the calibrated catalog improves more than just the fitted observables. However, they are still measured on the calibration volume. The absence of released parameter values compounds the issue: even the intended form of Eq. 2.5 is ambiguous (it appears to interpolate linking lengths but is written with Np notation), and without the fitted constants the claimed recipe cannot be independently applied. For a mock-generation paper aimed at DESI/LSST, an external transfer test is the decisive missing element. The right verdict remains CONDITIONAL: accept the method as a promising calibrated scheme, but require one independent validation before the headline claim is taken at face value. No change to the reader's verdict is needed.","tokens_in":13368,"tokens_out":6112,"duration_ms":61832,"concrete_test":"Run an independent reference simulation with a different cosmology, box size, and random seed (e.g., TNG100-2-Dark or an OuterRim/Quijote dark-matter run), together with its matched lower-resolution FastPM realization using the same initial linear density field and seed. Apply the published relaxed-FOF l(Np,i,z), r0(z), and PGD parameters without refitting, then measure the ratios of halo mass function, linear halo bias, and matter P(k) to that reference at z=0, 0.5, 1, and at k up to 10 h/Mpc. If the ratios leave the 3% bias band and roughly 10% mass-function band claimed in Figs. 3–4, transfer fails. As a second check, build a full N-body lightcone from the reference and compare the PGD-corrected C_l against it instead of halofit; halofit is not a sufficient standard at the claimed few-percent lensing accuracy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that calibrated FastPM is comparable to high-resolution N-body at the same mass resolution with two orders fewer steps—rests on calibration functions transferring to new simulations. The linking lengths l(Np,i,z), the fake-halo threshold r0(z) (Eqs. 2.2–2.6), and the PGD parameters {alpha0, A, B, kl,0, gamma, ks,0} (Eqs. 3.5–3.6) are fitted to TNG300-2-Dark's mass function, halo bias, and matter P(k) in a 205 Mpc box, and every validation figure uses that same box with the same initial linear density field. Because the tests reuse the calibration field, agreement at the low-mass end is partly by construction; the scatter among TNG300-1/2/3-Dark in Figs. 3–5 is a resolution/numerics comparison at fixed realization, not a statistical sample. Realization-specific calibration is a real risk: r0(z) is tuned to remove fake halos in high-density regions of this particular density field, and the PGD parameters are matched to P(k) over a limited k range of one box. The lightcone test (Fig. 13) also validates against halofit theory rather than against a full N-body lightcone, so it does not independently establish matter-field accuracy. The paper states the functions are 'simple functions we choose to produce correct halo mass function and halo bias' (Sec. 2.1), and no fitted values, code, or external validation are provided; therefore the headline accuracy is conditional on an out-of-sample test that is not reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript describes two modifications to the FastPM quasi-N-body code aimed at improving low-resolution mock catalogs. First, a 'relaxed-FoF' halo finder uses halo-mass- and redshift-dependent linking lengths plus a velocity-dispersion-based fake-halo rejection, with parameters calibrated so that the FastPM halo mass function and bias match TNG300-2-Dark. Second, a potential gradient descent (PGD) correction is inserted after each FastPM time step, with redshift-dependent parameters fitted to the TNG300 matter power spectrum. The authors compare halo bias, mass function, auto/cross power spectra, redshift-space power spectra, and catalog matching against TNG300-1/2/3-Dark, and show that a PGD-enabled FastPM lightcone improves the lensing convergence tomographic power spectrum relative to halofit. The concluding claim is that calibrated FastPM is comparable to a high-resolution N-body simulation at the same mass resolution with two orders of magnitude fewer time steps.","tokens_in":13693,"tokens_out":5548,"duration_ms":55338,"significance":"If the calibration functions transfer beyond the single 205 h^-1 Mpc box and the TNG300-2-Dark initial density field, the paper provides a practical recipe for producing large-volume mocks for DESI/LSST at substantially reduced cost. The paper's strengths are its direct comparisons with several TNG resolutions sharing the same initial field, its frank discussion of the resolution-limited behavior of small halos, and the embedding of PGD into the time integration so that lightcone outputs are consistently corrected. However, because the central statistics that define the calibration are matched by construction and no out-of-sample test is reported, the headline claim is currently conditional on a transferability assumption that the paper does not establish.","major_comments":[{"comment":"The relaxed-FoF parameters are introduced as 'simple functions we choose to produce correct halo mass function and halo bias.' Consequently, the agreement shown for the mass function (Fig. 4) and halo bias (Fig. 3) is a statement about the calibration target, not an independent test. The paper should either provide an out-of-sample validation (e.g., a different realization, box size, or cosmology, or at least a split-sample procedure within the same box) or explicitly reword the claims so that these halo comparisons are presented as calibration checks. The abstract's 'comparable to high resolution N-body' claim needs this evidence.","section":"Sec. 2.1 (Eqs. 2.2–2.6) and Figs. 3–4"},{"comment":"The PGD parameters are fitted to match the matter power spectrum of the reference simulation, so the improved P(k) in Fig. 11 is by construction at the fitted redshifts. The out-of-sample lightcone test in Fig. 13 compares the measured convergence power spectrum to the halofit analytic prediction rather than to a full N-body lightcone, and no error bars or significance levels are given; this does not by itself validate the matter field accuracy for lensing. Please add a comparison to a full N-body (or at least a validated emulator) lightcone, or soften the lensing claim accordingly.","section":"Sec. 3.1 (Eqs. 3.5–3.6) and Fig. 13"},{"comment":"The fitted values of the free parameters l1, l6, A1, A2, B1, B2, alpha0, A, B, kl,0, gamma, and ks,0 are never reported. Without these values, the 'calibration' cannot be reproduced or tested, and the reader cannot assess whether the functional forms are stable or overfit. A table of the best-fit parameters and, ideally, a description of the fitting procedure and the covariance of the fits should be included.","section":"Secs. 2.1 and 3.1"},{"comment":"The text says the mass resolution of the FastPM run used for the PGD fit is 125 times lower than TNG300-2-Dark, which would correspond to 250^3 particles, but the halo section uses 625^3 particles (8 times lower). The particle number and box for the PGD calibration runs are not stated; please clarify this inconsistency and specify which simulations produced Fig. 11.","section":"Sec. 3.1"}],"minor_comments":[{"comment":"In line 18 of the algorithm, the loop variable is declared as i but the body references halo[j]; this appears to be a typo for halo[i].","section":"Algorithm 1"},{"comment":"The filter scales k_l and k_s are introduced via k^2/k_l^2 and k^4/k_s^4, but their units and the fitted ranges are not stated; please define them and report the fitted values.","section":"Eqs. (3.3)–(3.4)"},{"comment":"The labels 'FastPM + PGD fit TNGDark' and 'FastPM + PGD fit TNG' are not explained in the body; specify which reference simulation was used for each fit and whether the same functional form is used for both.","section":"Fig. 11 caption and Sec. 3.1"},{"comment":"The phrase 'The shadow region shows the 1% deviation' should read 'shaded region'; the same figure captions should consistently use 'shaded'.","section":"Page 15"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a useful technical demonstration, but the central claim of survey-scale utility depends on calibration transfer that is not tested. The missing parameter table and the inconsistency about the PGD calibration resolution should be fixed in revision. I would ask for at least one out-of-sample validation (different seed, box, or cosmology) before considering this ready for publication; without it, the paper is better framed as a proof-of-concept calibration than as a validated mock-generation recipe."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid methods contribution with one genuinely new trick, but the headline accuracy claim is only established for a single simulation; the transfer to DESI/LSST mocks is an untested leap.\n\nThe new piece is relaxed-FOF: a linking length that grows for low-mass halos and with redshift, plus a velocity-dispersion cut to reject fake halos. That is not in the cited literature and it does what it says: the missed-halo fraction and halo mass function at M ~ 1e11 Msun improve a lot, and several independent statistics (halo auto power, RSD, halo-matter cross) also improve. Embedding PGD per time step, rather than as a post-processing fix, is also a nice extension of the authors' earlier work and gives consistent lightcone outputs. The paper is honest about imperfect spots, e.g. low-z small-scale halo power overprediction.\n\nThe soft spot is the one the stress-test note highlights: the calibration functions are chosen to reproduce the reference mass function, bias, and matter P(k) on TNG300-2-Dark, and all validation figures are on the same box with the same initial density field. So the excellent agreement on those quantities is partly by construction. The independent checks are not fully independent because they use the same simulation realization; there is no test on a different cosmology, box size, or even a different random seed. The lightcone lensing test compares against halofit, not a full N-body lightcone. No code, data, or fitted parameter values are released, so other groups cannot test transfer without reimplementing.\n\nThat said, the paper does not hide the fitting—it says the functions are 'simple functions we choose.' The limitation is an omitted out-of-sample test, not a load-bearing mathematical error. The method is plausible and the comparisons are thorough. I would send it to peer review, with a clear request for an external validation experiment; even one additional simulation with a different seed or box would substantially raise confidence.","headline":"Useful methods paper with a genuinely new halo finder, but the accuracy claim rests on a single calibration simulation and needs an out-of-sample test before it can be trusted for survey mocks.","tokens_in":14280,"tokens_out":1990,"would_cite":true,"duration_ms":19866,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.80.-k","95.35.+d","98.65.-r"],"model":"deepseek-v4-flash","headline":"Calibrated low-resolution FastPM simulations reproduce full N-body halo and matter statistics at roughly one hundredth of the time steps.","keywords":["fast simulations","quasi-N-body","halo finder","friends-of-friends","potential gradient descent","halo bias","matter power spectrum","weak lensing convergence"],"falsifier":"Run FastPM with the published relaxed-FoF and PGD parameters on an independent simulation with a different cosmology or box size, and compare the halo mass function, halo bias, and matter power spectrum against a matching full N-body simulation; deviations larger than the few-percent agreement shown for the calibration run would show that the calibration does not transfer.","tokens_in":13075,"feed_emoji":"🌌","tokens_out":9655,"duration_ms":89977,"temperature":0.7,"pith_summary":"This paper argues that fast, low-resolution particle-mesh simulations can be calibrated to produce halo and matter statistics that rival high-resolution full N-body simulations of the same mass resolution, at about two orders of magnitude fewer time steps. For halos, the key is a modified friends-of-friends finder, relaxed-FoF, whose linking length grows for smaller and higher-redshift halos and which rejects spurious, unbound clumps by their velocity dispersion. For the matter field, the paper embeds a potential gradient descent (PGD) correction into every time step, adding a sub-grid displacement that restores power on nonlinear scales. The result is a FastPM run that matches high-resolution reference halo bias, mass function, real- and redshift-space clustering, halo-matter cross-power, and weak lensing convergence power spectra closely enough for survey-scale mock catalogs.","feed_headline":"Calibrated low-res FastPM matches N-body stats with 100x fewer steps","feed_subtitle":"A mass-dependent halo finder and per-step subgrid corrections close the gap to high-resolution runs.","key_machinery":"Two calibrated corrections carry the argument. The first is relaxed-FoF, a halo finder that makes the linking length a function of halo particle number and redshift, $l(N_{p,i},z)$, and filters out fake halos using a velocity-dispersion threshold $r_0(z)$ relative to the expected mass–velocity-dispersion scaling; this reassembles fragmented low-resolution halos and removes unbound false detections. The second is potential gradient descent (PGD) embedded into every FastPM time step, which adds a particle displacement along the gradient of a filtered gravitational potential, with parameters $\\alpha$, $k_l$, and $k_s$ chosen as simple functions of scale factor; this restores small-scale matter power and makes static snapshots and lightcone outputs consistently corrected.","core_discovery":"On the halo side, the paper claims that the dominant failure of low-resolution FastPM is not force resolution but halo identification: small halos fragment under the standard linking length $l=0.2$, producing missing and overly clustered objects, while dense regions generate fake halos. Relaxed-FoF uses linking lengths $l(N_{p,i},z)$ that increase for smaller halos and at higher redshift, defined across six particle-number bins with linear interpolation (Eqs. 2.2–2.5), and removes candidates whose velocity dispersion exceeds $r_0(z)$ times the expected dispersion from the scaling relation (Eqs. 2.1, 2.6). With initial conditions generated on a mesh twice as fine as the particle grid, the missed-halo fraction at $10^{11}\\,M_\\odot$ drops below 10% at $z=2$, and halo bias, mass function, auto and redshift-space power spectra, cross-correlation with the reference catalog, and halo-matter cross-power all improve. On the matter side, embedding PGD after each FastPM step, with redshift-dependent parameters $\\log(\\alpha/\\alpha_0)=Aa^2-Ba$ and $k_l=k_{l,0}a^\\gamma$, restores the nonlinear matter power spectrum to roughly 1% agreement at fixed redshifts. A lightcone built by interpolating particle positions between steps then yields weak lensing convergence tomographic power spectra that agree substantially better with the theoretical nonlinear model.","pith_inferences":["The calibration functions—linking length bins, fake-halo threshold, and PGD parameters—are fitted to a single reference simulation, and the paper does not demonstrate that they transfer to another cosmology, box size, or initial density field; testing transferability on an independent simulation is the natural next step.","Relaxed-FoF changes how mass is assigned to individual halos, so statistics that depend on internal structure, such as concentration or assembly bias, may shift in ways not captured by the power-spectrum tests presented.","Because PGD alters particle positions in every step, it also changes particle velocities; the method's small-scale velocity statistics beyond the tested redshift-space halo power remain a useful further check.","A direct extension would be to fit the same functional forms to several independent reference simulations and ask whether the fitted parameters drift; stable parameters would strengthen confidence in the method's transferability."],"forward_implications":["A low-resolution FastPM run with relaxed-FoF produces halo catalogs whose missed-halo fraction, bias, mass function, and real- and redshift-space power spectra are comparable to those of a full N-body run at the same particle mass.","For halos above $3\\times10^{11}\\,M_\\odot$, the calibrated simulation keeps deviations in redshift-space halo power and halo-matter cross-power within about 6% out to $k=2\\,h\\,\\mathrm{Mpc}^{-1}$ and $z\\le2$.","Embedding PGD at every time step brings static-snapshot matter power to roughly 1% agreement with the reference and makes lightcone outputs consistently corrected, so weak lensing convergence tomographic power spectra approach the theoretical prediction.","Because the corrected FastPM needs only about 20 time steps for a lightcone, survey-scale mock generation over cubic-gigaparsec volumes becomes feasible at the $10^{11}\\,M_\\odot$ halo resolution required by future surveys.","The relaxed-FoF catalog also has better large-scale auto power in real and redshift space, and better halo-matter cross-power, than a full N-body simulation at the same mass resolution."],"supporting_citations":[{"why":"Introduces the FastPM scheme and provides the standard friends-of-friends baseline that relaxed-FoF improves on.","marker":"[4]"},{"why":"Introduces the potential gradient descent displacement model that this paper embeds into every FastPM step.","marker":"[10]"},{"why":"Supplies the high-resolution reference simulation and companion runs used for calibration and comparison.","marker":"[11]"},{"why":"Provides the velocity dispersion–mass scaling relation used to identify and remove fake halos.","marker":"[12]"},{"why":"Computes the power spectra used throughout the halo and matter comparisons.","marker":"[13]"},{"why":"Provides the halofit nonlinear matter power spectrum used as the theoretical target for the weak lensing convergence spectra.","marker":"[18]"},{"why":"Motivates the mass-resolution target by quantifying systematics for halos with a few hundred particles.","marker":"[9]"}],"fun_headline_variants":["Relaxed halo finder brings FastPM to N-body accuracy","Low-res FastPM hits N-body stats via relaxed-FoF and PGD","100x faster fast simulations reach N-body halo accuracy","Relaxed-FoF and PGD make FastPM rival high-res N-body","New halo finder and subgrid fixes give FastPM N-body accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The linking lengths, fake-halo threshold, and PGD parameters are all fitted to one reference simulation, and the paper never tests whether those fixed functions remain accurate for different cosmologies, box sizes, or initial density fields before recommending them for survey mocks.","fun_headline_variants_meta":{"raw":{"variants":["Relaxed halo finder brings FastPM to N-body accuracy","Low-res FastPM hits N-body stats via relaxed-FoF and PGD","100x faster fast simulations reach N-body halo accuracy","Relaxed-FoF and PGD make FastPM rival high-res N-body","New halo finder and subgrid fixes give FastPM N-body accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000646,"raw_usage":{"total_tokens":3033,"prompt_tokens":1078,"completion_tokens":1955,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":694,"completion_tokens_details":{"reasoning_tokens":1859}},"tokens_in":694,"tokens_out":1955,"duration_ms":15225,"temperature":1.0,"reasoning_tokens":1859,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:18:27.668653+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run FastPM with the published relaxed-FoF and PGD parameters on an independent simulation with a different cosmology or box size, and compare the halo mass function, halo bias, and matter power spectrum against a matching full N-body simulation; deviations larger than the few-percent agreement shown for the calibration run would show that the calibration does not transfer.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Computes the power spectra used throughout the halo and matter comparisons."}],"review_version":1}