{"id":"f304afc8-07ca-43ba-b6f0-9835410a0f8f","arxiv_id":"2602.03382","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Adding peculiar-velocity clustering to galaxy clustering on non-linear scales tightens forecast fσ8 constraints from 4.7% to 3.8% in realistic mocks, but biases the recovered HOD parameters.","lead":"The paper trains machine-learning emulators on N-body simulations to model how galaxies and their peculiar velocities cluster together on small, non-linear scales. It finds that adding velocity information improves the predicted precision on cosmic growth from about 4.7% to 3.8%, while mock tests show biases in galaxy-formation parameters.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Realistic-mock validation is self-consistent by construction: mocks and emulator share the same HOD, so the claimed unbiased fσ8 recovery and 3.8% gain are untested against alternative galaxy-halo relations.","rationale":"The paper's central claim is precisely the unbiasedness of cosmological parameters and the precision gain. The weakest point is the lack of external validation of the galaxy-velocity connection. All training and validation sets come from HOD on AbacusSummit; the realistic mocks add observational noise but not model misspecification. The authors explicitly flag this in Sec. 5.2. Thus, the 'unbiased' result and the 3.8% figure are established only within a closed loop. A test against an independent galaxy-halo model is the minimal check that would settle the concern. The reader identified the same assumption, so I agree. Since the reader's verdict is already CONDITIONAL, my read does not change it.","tokens_in":22527,"tokens_out":3766,"duration_ms":40438,"concrete_test":"Generate realistic mocks with the same halo catalogue but a different galaxy-formation prescription—e.g., subhalo abundance matching or a velocity-biased HOD (satellite velocities scaled by 1.1 or with an off-center velocity distribution)—and repeat the Sec. 5.1 MCMC fits with the same emulator and data vector. If the recovered fσ8 shifts by more than ~0.5σ or the joint analysis no longer beats galaxy-only (4.7% degraded), the claimed unbiased improvement is an artifact of self-consistency. Alternatively, use a hydrodynamical simulation galaxy catalogue (e.g., IllustrisTNG or Magneticum) at similar number density to check the same.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the central claim that joint galaxy+velocity clustering on non-linear scales improves fσ8 constraints and remains unbiased, the paper must show that the HOD emulator can marginalise over the actual systematic differences between mock and real data. The validation in Sec. 5 does not do this: the realistic mocks are generated from the same AbacusSummit halos, same HOD formalism, and same halo-mass definition used to train the emulator. The paper states this directly: 'our noisy sample is built from the same underlying formalism' (Sec. 5.2). Consequently, the HOD parameters absorbing the added noise is an interpolation within the training model class, not evidence of robustness to unmodelled physics such as velocity bias, assembly bias, or correlated distance errors. The treatment of the distance-indicator conversion is a second self-consistency: 'we ... use the true underlying cosmology (Planck18) to compute velocities from log-distance ratios' (Sec. 5.1), so the mock analysis removes a systematic that would be present in real data when the fiducial cosmology is imperfect. If a real survey's galaxy-velocity relation differs from the HOD model, the same flexibility that absorbs noise in these mocks could absorb signal and shift cosmological parameters; the 'unbiased' recovery and the 3.8%-vs-4.7% improvement are therefore conditional on the HOD being the correct generative model. The poor reduced chi-squared for vv (4.3) and vg (3.0) in Table 3 already hints that the model is not a good description of the noisy velocity statistics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a simulation-based emulator for the redshift-space two-point statistics of galaxies and peculiar velocities on non-linear scales (0.3–60 h⁻¹ Mpc): the galaxy autocorrelation monopole/quadrupole, the velocity autocorrelation monopole/quadrupole, and the galaxy–velocity cross-correlation dipole. Using 88 AbacusSummit cosmologies × 600 HOD models at z = 0.2, the authors build five multi-scale Gaussian-process emulators with a Kronecker-structured kernel and propagate both cosmic variance and emulator error into the likelihood. On a recovery set of 25 Planck2018 full-sky mocks they report unbiased cosmological and HOD parameters, with the joint analysis improving fσ₈ precision from 1.5% (galaxy clustering only) to 1.1%. They then construct realistic mocks reproducing DESI BGS galaxy densities, ZTF SNe + TF/FP velocity-tracer densities, distance-indicator scatter, FKP weights, and Alcock–Paczynski distortions. Fitting these mocks yields biased HOD parameters but, the paper claims, unbiased cosmological parameters, with a joint fσ₈ precision of 3.8% versus 4.7% from galaxy clustering alone. The abstract concludes that velocity statistics add information on non-linear scales, but with diminished returns under realistic measurement conditions.","tokens_in":22971,"tokens_out":22738,"duration_ms":217546,"significance":"The central claim — that galaxy+velocity clustering on non-linear scales improves fσ₈ constraints (3.8% vs 4.7% in realistic mocks) while remaining cosmologically unbiased — is, if established, a useful result for low-redshift peculiar-velocity surveys and for the design of future samples. The paper has real methodological strengths: the emulator is validated against withheld cosmologies/HODs and on 25 independent realisations; the code (MKGpy) and data are public or promised on Zenodo; the covariance treatment, including emulator error and the Percival et al. (2022) precision-matrix correction, is careful; and the limitations are disclosed unusually candidly (§5.1, §5.2). However, the realistic-mock demonstration is a self-consistency test: the mocks share the same HOD formalism as the emulator, the η→v conversion uses the true cosmology, the velocity statistics fit the data poorly (rχ² = 4.3 and 3.0), and the joint fσ₈ estimate shows a systematic shift of ≈3.7σ of the mean. The headline claims are therefore conditional on assumptions that the current validation does not yet test.","major_comments":[{"comment":"For the realistic mocks, the velocity-only and galaxy-velocity likelihoods are rejected by the data: rχ² = 4.3 (p ≈ 2×10⁻²⁰) for vv and 3.0 (p ≈ 6×10⁻⁵) for vg. The Gaussian likelihood (Eq. 14) is therefore not a valid description of these measurements, and the per-realisation uncertainties and unbiased-recovery claims from any analysis containing vv or vg are not well grounded. The joint analysis instead gives rχ² = 0.4 (p ≈ 1−10⁻¹⁰), which the authors attribute to overfitting by the HOD parameters (Sec. 5.2): the velocity misfit is masked by parameter flexibility, not explained. Please calibrate the velocity noise model (e.g., correlated or non-Gaussian η errors after FKP weighting, HOD-dependence of the covariance, extra small-scale variance) so that the vv and vg fits are acceptable, and re-check the joint conclusions with the calibrated likelihood.","section":"Sec. 5.1, Table 3"},{"comment":"The realistic-mock analysis converts log-distance ratios to peculiar velocities using the true Planck18 cosmology while applying the AP distortion with the c003 fiducial cosmology; the text defers study of this effect. In real data the same fiducial enters both conversions via D(z) and H(z) in Eqs. (B.4)–(B.5), so using the truth removes a systematic that any analysis will face. Since the headline 'unbiased cosmological recovery' and the 3.8% fσ₈ figure come from these mocks, this is load-bearing. Please re-run the inference (even on a subset of the 25 mocks) with the η→v conversion computed with the c003 fiducial, and report the effect on the fσ₈ bias, the HOD biases, and the joint gain.","section":"Sec. 5.1, App. B Eqs. (B.4)–(B.5)"},{"comment":"The paper states 'our noisy sample is built from the same underlying formalism' as the emulator, and Sec. 6 concludes that 'the HOD model is flexible enough to marginalise over the noisy clustering without biasing the cosmological parameters.' Because the realistic mocks use the same AbacusSummit halos, the same Zheng et al. HOD, and the same training ranges, the validation is an interpolation within the training model class; it does not test robustness to real galaxy–halo connection differences (velocity bias, assembly bias, correlated distance errors). This is the main untested assumption behind the claim of unbiased cosmology in realistic conditions. Please either add a stress test with a galaxy model outside the HOD class (e.g., an HOD with velocity bias, or subhalo abundance matching), or reword the abstract and conclusions to state that unbiasedness and the 3.8% gain are demonstrat","section":"Sec. 5.2"},{"comment":"In the realistic joint ('tot') analysis, ⟨Δfσ₈⟩ = 1.32×10⁻² with ⟨σ_θ⟩ = 1.78×10⁻². Averaged over 25 realisations the error of the mean is ≈0.36×10⁻², so the shift is ≈3.7 times the expected dispersion of the mean. The text itself concludes this 'point[s] toward a potential systematic bias that could cancel out the statistical gain.' The claimed gain over galaxy-only (0.38×10⁻² reduction in σ) is small compared with this shift, and the abstract's 'cosmological constraints remain unbiased' is stronger than the evidence. Please provide a formal significance test of the fσ₈ shift over the 25 realisations and report the bias-variance budget explicitly.","section":"Sec. 5.1, Table 3 (fσ₈ row)"},{"comment":"The emulator covariance is fixed at the true simulation parameters during inference. In the realistic mocks the HOD MAP moves far from the truth (Table 3: log M₁ biased by ≈7σ, α by ≈2.3σ), so the fixed covariance is not the covariance at the fitted point; this affects both the error bars and the log|C_tot| term in Eq. (14). In addition, σ(Z_θ) < 1 for every parameter in every analysis of both Tables 2 and 3; the authors note this could indicate overestimated uncertainties. With inflated error bars, a real bias (e.g., the fσ₈ shift noted above) is masked. Please report the per-realisation rχ² distribution for each analysis and, for at least a few realisations, rerun the chains evaluating C_tot at the MAP or at each step to check that the quoted 1σ errors are stable.","section":"Sec. 4.1, Tables 2–3"}],"minor_comments":[{"comment":"The S values appear to contain a typo: '10⁵ for ξ_gg²' should presumably be '10⁵ for ξ_vv²', since ξ_gg² was already assigned 10¹ and ξ_vv² is the remaining multipole.","section":"Sec. 3.1, Eq. (13)"},{"comment":"The paper states it exclusively uses z = 0.2 snapshots (Sec. 2.1), but the recovery and realistic mocks are described with a radial cut 'corresponding to a redshift range of z ∈ [0, 0.1]' (Sec. 2.3, Fig. 1) and densities matching n(z) within z ∈ [0, 0.1] (Sec. 2.4). Please clarify how a z = 0.2 snapshot is mapped to a z ≤ 0.1 survey; this also bears on the AP and η→v conversions in Sec. 5.1 and on the quoted fσ₈(z = 0.2).","section":"Sec. 2.1 vs Secs. 2.3–2.4"},{"comment":"Typo: 'The observables we which to emulate' should read 'we wish to emulate'.","section":"Sec. 2.3"},{"comment":"Typo: 'In the following test we our models' should read 'we test our models'.","section":"Sec. 3.2"},{"comment":"The p-value notation '1−1×10⁻¹⁰' is confusing; write '> 1−10⁻¹⁰' or '≈ 1'. Likewise, 'poor p-values, larger than 1−1×10⁻³' (Sec. 5.2) should be rephrased: p-values near unity indicate overfitting, not a poor fit in the usual sense.","section":"Sec. 5.1, Table 3; Sec. 5.2"},{"comment":"The units of P_v are given as 'h̄³ Mpc³ km² s⁻²'; the 'h̄' appears to be a typo for 'h⁻³'. Please also define P_g and P_v in one place and check their dimensions.","section":"Sec. 2.4, Eq. (11)"},{"comment":"Notation is inconsistent: Fig. 2's caption and Fig. 3's axis label use 'ξ_gv', while the text and Eq. (3) use 'ξ_vg'. Harmonize throughout.","section":"Figs. 2 and 3; text"},{"comment":"Harmonize the definition of η with Sec. 2.4 (D(z_obs)/D_obs vs D(z_cos)/D(z_obs)) and state the validity range of the low-redshift linearization (at z = 0.2, for v up to ~1000 km/s), since the realistic-mock conversion relies on it.","section":"App. B, Eqs. (B.4)–(B.5)"}],"recommendation":"major_revision","confidential_remarks":"This is an honest, well-structured methods paper; the authors disclose the main limitations in the text, and I see no indication of fabrication or citation problems. The principal risk is that the realistic-mock validation is a self-consistency test: the poor vv/vg fit quality (rχ² = 4.3, 3.0), the use of the true cosmology for η→v conversion, and the ≈3.7σ-of-the-mean fσ₈ shift in the joint fit are the points that must be addressed before the '3.8% vs 4.7%' result is used as evidence of readiness for real data. My major_revision recommendation rests on four requested items: (a) a recalibrated velocity likelihood with acceptable vv/vg fits, (b) a test of the η→v fiducial-cosmology dependence, (c) at least one robustness test against a non-HOD galaxy model, and (d) a quantitative bias test for fσ₈. If these are delivered, the paper would be a solid A&A contribution; the emulator framework itself (Kronecker GP, public code, careful covariance treatment) is a genuine value-add."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a method-and-forecast paper, and within that scope it is solid. The authors extend their existing multi-scale Gaussian-process emulator to non-linear galaxy-galaxy, velocity-velocity, and galaxy-velocity multipoles, then run a realistic mock campaign with SNIa and TF/FP velocity errors matching ZTF+DESI densities. The headline result — joint fitting gives 3.8% precision on fσ8 versus 4.7% from galaxies alone, with unbiased cosmological parameters — is reproduced in their mocks and is a genuinely new quantitative statement.\n\nWhat it does well: the machinery is inherited from Dumerchat & Bautista (2024), but the application to vv and vg on non-linear scales is new, and the execution is careful. The estimator derivation in Appendix A is thorough. They report biases via FoB and constraining power via FoM rather than just showing contours. The covariance rescaling follows Percival et al. (2022). And the paper is transparent: data on Zenodo, code public, and the limitations are stated in the text, not buried.\n\nThe soft spots are real but mostly admitted. The key one is that the realistic mocks are built from the same HOD formalism used to train the emulator — the paper says this directly in Sec. 5.2. So the HOD parameters absorbing the velocity noise is interpolation inside the training model class, not evidence of robustness to velocity bias, assembly bias, or correlated distance errors. Similarly, Sec. 5.1 says velocities are computed from log-distance ratios using the true Planck18 cosmology, removing a systematic that real data would contain. The poor reduced chi-squares for vv (4.3) and vg (3.0) with near-zero p-values in Table 3 are a further warning that the model does not actually fit the noisy velocity statistics well; the unbiased cosmology is likely coming from HOD flexibility overfitting the small-scale signal, which they also state. So the 3.8% gain is a conditional forecast, not a measurement claim.\n\nMinor issues: the AP treatment is first-order only, and the emulator covariance is fixed at the true parameters. Both are acknowledged and reasonable for this kind of study. The stress-test note is right on target, but the authors already flag it; the remaining gap is external validation against a different galaxy-formation model.\n\nThis paper deserves a serious referee. The main thing I would ask the authors to do is either run one validation against a non-HOD galaxy assignment (abundance matching or a hydro simulation) or explicitly downgrade the headline to \"conditional on the HOD being the generative model.\" For someone working on velocity surveys or emulator-based forecasting, this is worth a careful read and worth citing for the concrete precision numbers.\n\nRecommendation: send to peer review. It is honest, reproducible, and the central claim holds within its stated scope.","headline":"A genuinely useful and unusually honest emulator-forecast paper: the 3.8%-vs-4.7% gain in fσ8 is real inside their mocks, but the paper itself admits it is a self-consistency test, not a claim about real-data robustness.","tokens_in":23613,"tokens_out":1846,"would_cite":true,"duration_ms":20758,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Jointly fitting galaxy and peculiar-velocity clustering on non-linear scales improves cosmological constraints, reaching 3.8% precision on the growth-rate parameter fσ8 in realistic mocks—versus 4.7% from galaxy clustering alone.","keywords":["peculiar velocities","galaxy clustering","emulator","halo occupation distribution","non-linear scales","growth rate fσ8","two-point correlation functions","mock catalogues"],"falsifier":"Build galaxy mocks from a physically different galaxy-halo connection—e.g., abundance matching, assembly-bias HOD, or hydrodynamical simulations—then run the emulator inference; if the recovered cosmological parameters (especially fσ8) shift by more than the claimed uncertainty, the 3.8% precision claim does not transfer to real data. A cheaper check: re-run the realistic-mock fits converting log-distance ratios to velocities using the same fiducial cosmology as the AP distortion instead of the true cosmology; if fσ8 moves by a significant fraction of the 1.8% quoted gain, the result is fragil","tokens_in":22385,"feed_emoji":"🔭","tokens_out":6336,"duration_ms":139427,"temperature":0.7,"pith_summary":"The paper sets out to show that the galaxy–velocity cross-correlation, traditionally analysed only on large scales, still carries independent cosmological information on non-linear scales down to 0.3 h⁻¹ Mpc. Using an N-body simulation suite and the halo occupation distribution (HOD) formalism, the authors train Gaussian-process emulators for the galaxy-galaxy, velocity-velocity, and galaxy-velocity two-point correlation functions in redshift space, then fit them to mock data. In ideal mocks the joint fit tightens constraints on σ8 and w0 and reaches 1.1% precision on the growth-rate parameter fσ8, compared to 1.5% from galaxy clustering alone. In realistic mocks with sparse tracers and Gaussian distance errors typical of type-Ia supernovae and Tully-Fisher/fundamental-plane measurements, the gain shrinks but survives: fσ8 is measured to 3.8% instead of 4.7%, and cosmological parameters remain unbiased although the galaxy-formation (HOD) parameters become systematically biased. The paper's central claim is that velocity statistics add usable information on non-linear scales, but that extracting it requires controlling small-scale noise and systematics.","feed_headline":"Peculiar velocities sharpen growth-rate constraints to 3.8%","feed_subtitle":"Jointly fitting velocities and galaxy clustering beats galaxies alone in realistic mocks, even with noisy distances.","key_machinery":"The load-bearing machinery is a multi-scale Gaussian process emulator whose kernel factorises as the Kronecker product of kernels over cosmology, HOD parameters, and separation, exploiting the Kronecker-product structure of the training grid (88 cosmologies × 600 HOD models). It outputs mean predictions and an emulator covariance for the five correlation-function multipoles; the emulator covariance is added to a cosmic-variance covariance estimated from many small simulation boxes. The observables are measured with pair-count estimators using the flat-sky approximation, which the paper verifies against full-sky measurements. Galaxy assignment uses the standard HOD model with central and sate","core_discovery":"The paper claims that jointly modelling galaxy and peculiar-velocity clustering on non-linear scales produces tighter, unbiased cosmological constraints than galaxy clustering alone. The emulator predicts five redshift-space multipoles—the monopole and quadrupole of the galaxy and velocity auto-correlations plus the dipole of the galaxy-velocity cross-correlation—as a function of cosmological parameters, HOD parameters, and separation. When all five are fitted together on ideal full-sky mocks at z=0.2, all parameters are recovered within 1σ and the fσ8 precision improves from 1.5% to 1.1%. When the same inference is run on realistic mocks that include sparse tracer densities, Gaussian distan","pith_inferences":["If this transfers to real data, a ~3-4% growth-rate measurement at z≈0.2 would be unusually precise at low redshift, strengthening combined constraints on dark energy and modified gravity when paired with high-redshift probes.","The paper's own caveat points to the key test: apply the emulator to mocks with velocity bias, assembly bias, or galaxy populations from hydrodynamical simulations; if cosmological recovery becomes biased, the 3.8% claim will not survive contact with real galaxies.","A cheaper falsifying test is to vary the fiducial cosmology used to convert log-distance ratios to velocities (the paper fixes it to the true cosmology while applying Alcock-Paczynski shifts with a different one); if fσ8 moves by a meaningful fraction of the 1.8% gain, the improvement is partly an artifact of the analysis choice.","Separating supernova and Tully-Fisher/fundamental-plane velocity samples—different densities and errors—could recover more of the lost gain than the combined sample used here, and is a natural next step."],"forward_implications":["Combining galaxy and velocity clustering on non-linear scales tightens constraints on σ8 and w0 relative to either tracer alone, in both ideal and realistic mocks.","Realistic velocity measurement errors and sparse tracer densities reduce the gain but do not erase it: fσ8 precision improves from 4.7% to 3.8%.","The flexibility of the HOD model can absorb unmodelled small-scale noise, so cosmological parameters can stay unbiased even when HOD parameters are biased.","Angular scales as small as 0.33 h⁻¹ Mpc add cosmological information, so correcting small-scale observational systematics such as fibre collisions is worthwhile.","The velocity-only and cross-correlation-only analyses show larger systematic shifts on fσ8 once realistic noise is included, indicating the joint fit is needed to keep the measurement unbiased."],"fun_headline_variants":["Velocity and galaxy clustering jointly hit 3.8% growth precision","Combining velocities and galaxies tightens fσ8 to 3.8%","Non-linear emulator pairs velocity and galaxy clustering","Peculiar velocities sharpen growth-rate constraints to 3.8%","Joint velocity-galaxy fit beats galaxies alone on fσ8"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 3.8% result assumes the same HOD model used to build the mocks can absorb all unmodelled small-scale velocity noise without biasing cosmology, and that converting log-distance ratios to velocities with the true cosmology—while the Alcock-Paczynski shift uses a different fiducial cosmology—is harmless; if either gives way, the unbiased 3.8% measurement may not hold for real data.","fun_headline_variants_meta":{"raw":{"variants":["Velocity and galaxy clustering jointly hit 3.8% growth precision","Combining velocities and galaxies tightens fσ8 to 3.8%","Non-linear emulator pairs velocity and galaxy clustering","Peculiar velocities sharpen growth-rate constraints to 3.8%","Joint velocity-galaxy fit beats galaxies alone on fσ8"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000375,"raw_usage":{"total_tokens":1831,"prompt_tokens":733,"completion_tokens":1098,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":1021}},"tokens_in":477,"tokens_out":1098,"duration_ms":9880,"temperature":1.0,"reasoning_tokens":1021,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T04:58:45.711082+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build galaxy mocks from a physically different galaxy-halo connection—e.g., abundance matching, assembly-bias HOD, or hydrodynamical simulations—then run the emulator inference; if the recovered cosmological parameters (especially fσ8) shift by more than the claimed uncertainty, the 3.8% precision claim does not transfer to real data. A cheaper check: re-run the realistic-mock fits converting log-distance ratios to velocities using the same fiducial cosmology as the AP distortion instead of the true cosmology; if fσ8 moves by a significant fraction of the 1.8% quoted gain, the result is fragil","supporting_citations":[],"review_version":1}