{"id":"792e1fdd-87ca-4db5-a0d9-f6dd7c9264ef","arxiv_id":"2412.01888","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Red and blue simulated galaxies can be described by the same field-level effective field theory as halo mocks, but blue galaxies require the dedicated HMQ model rather than LRG-style HOD models.","lead":"This paper extracts the large-scale clustering parameters of red and blue galaxies from two hydrodynamic simulations and checks them against cheaper 'painted galaxy' mock catalogs. The match holds for red galaxies and for blue galaxies only when a dedicated model is used, which matters for how future DESI data will be interpreted.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ELG vs LRG-HOD tension is tested at mismatched redshifts: LRG-HOD priors are at z=0.5 while hydro ELGs are at z=1, so the central claim 3 lacks a same-epoch comparison.","rationale":"The reader's weakest_assumption focused on the sSFR selection proxy for observed LRGs/ELGs. That is a legitimate external-validity concern, but the paper's internal comparison between hydro galaxies and HOD/HMQ mocks uses the same selection cuts, so it is not the most load-bearing issue for the central claim. The more immediate threat is the redshift mismatch in the ELG vs. LRG-HOD comparison: the paper asserts strong tension and even 'no decorated HOD model' can match ELGs, yet the LRG-HOD priors are computed at z=0.5 while the ELG measurements are at z=1. Since bias parameters evolve with redshift and the halo population at fixed HOD parameters changes, the comparison is not apples-to-apples. The paper's HMQ comparison uses z=1.1, which is closer but still not exactly z=1. The paper's analytic argument (Section 4.3) is explicitly simplistic and ignores satellites, so it does not close the gap. The proposed test is straightforward because the same AbacusSummit small suite and the same EFT pipeline are already in place; generating mocks at z=1 would settle whether the tension persists in a same-epoch comparison. Until that check is done, the strong wording of item 3 is not fully supported, hence the recommendation to accept conditionally. The paper's other strengths — field-level EFT measurements from two large hydro simulations, the confirmation of LRG-HOD consistency for LRGs at matched redshift, the HMQ agreement, and the normalizing-flow inversion — are well executed and support the broader conclusions, so no rejection is warranted.","tokens_in":24640,"tokens_out":16231,"duration_ms":177236,"concrete_test":"Generate a set of LRG-HOD and HMQ mock catalogs from the AbacusSummit small suite at z=1 (rather than z=0.5 for LRG-HOD and z=1.1 for HMQ), run the same field-level EFT pipeline used for the hydro simulations, and compare the resulting EFT parameter distributions with the MTNG and Astrid ELG measurements. If the ELG data points fall inside the LRG-HOD contours at z=1, the claimed tension is not supported; if they remain outside while lying within HMQ contours, the conclusion is robust.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's key novel claim (Section 2, item 3) is that ELG local bias parameters are in very strong tension with LRG-based HOD models, and that no decorated HOD can reproduce them. The evidence, however, compares MTNG/Astrid ELG measurements at z=1 (Section 3.2) against LRG-HOD prior predictive distributions explicitly generated at baseline redshift z=0.5 (Section 3.3: \"Our baseline redshift for both LRG HOD catalogs is z = 0.5\"). The bias parameters b1, b2, b3 are redshift-dependent: the Eulerian bias evolves as b(z) = 1 + (b_L - 1)/D(z), and the halo mass function at fixed HOD parameters changes between z=0.5 and z=1. Therefore, the prior distribution of EFT parameters for the same HOD parameter priors need not be the same at z=1. The paper does not generate LRG-HOD mocks at z=1 for this comparison; it instead offers a toy analytic model (Section 4.3) and a comparison to HMQ mocks at z=1.1, which is still not exactly z=1. If the LRG-HOD prior distribution at z=1 shifts toward the ELG measurements, the central statement that no LRG-like HOD can describe ELGs would be weakened. This is a concrete, testable gap in the evidence for one of the paper's main conclusions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper measures the effective field theory (EFT) bias parameters and redshift-space counterterms for LRG and ELG samples selected from the MillenniumTNG (MTNG) and Astrid hydrodynamic simulations, and compares them with the corresponding quantities extracted from HOD (LRG) and HMQ (ELG) mock catalogs at the field level. The authors report that LRG EFT parameters are consistent with previous LRG-HOD priors, that ELG local bias parameters are in tension with LRG-like HOD models but consistent with the HMQ model, that ELGs show weaker fingers-of-God, that the response to sSFR cuts differs between MTNG and Astrid owing to feedback modeling, and finally that normalizing-flow-based conditional distributions can map measured EFT parameters to optimal HOD/HMQ parameters.","tokens_in":24993,"tokens_out":6148,"duration_ms":62707,"significance":"If the results hold, the paper provides a valuable field-level validation of halo-based galaxy-halo connection models for both red and blue galaxies, directly supporting the use of HOD/HMQ-based priors in EFT full-shape cosmological analyses. The comparison across two independent hydrodynamic simulations with different feedback implementations is a genuine strength, as is the use of the publicly available Hi-Fi mocks code and the explicit field-level cross-correlation analysis. The finding that ELG local bias parameters deviate from LRG-HOD predictions, if confirmed at matched redshift, would have practical implications for ELG analyses.","major_comments":[{"comment":"The central claim that ELG local bias parameters are in 'very strong tension' with LRG-based HOD models compares MTNG/Astrid ELG measurements at z=1 (Section 3.2) with LRG-HOD prior distributions generated at baseline redshift z=0.5 (explicitly stated in Section 3.3). Bias parameters evolve with redshift and the halo mass function at fixed HOD parameters changes between z=0.5 and z=1, so the prior distribution of EFT parameters for the same HOD priors need not be the same at z=1. The HMQ baseline z=1.1 is also not exactly z=1. Please add a same-epoch comparison, for example by generating LRG-HOD mocks at z=1 or by mapping the z=0.5 prior distributions to z=1 using the growth factor and mass function, and demonstrate that the conclusion is unchanged.","section":"Section 3.3 and Section 2, item 3"},{"comment":"Astrid redshift-space EFT fits use kmax=0.4 h/Mpc in both real and redshift space, whereas MTNG and the HOD/HMQ mocks use kmax=0.2 h/Mpc for redshift-space transfer function fits. Since redshift-space counterterms are known to be sensitive to the scale cut, the reported consistency of Astrid counterterms with HOD/HMQ priors in figs. 7 and 10 could be affected by the different cutoff. Please provide Astrid redshift-space results at kmax=0.2 h/Mpc or explicitly demonstrate that the counterterm constraints are stable when kmax is varied between 0.2 and 0.4 h/Mpc.","section":"Section 3.4"},{"comment":"The analytic argument explaining the enhanced b2 of LRGs relative to halos is derived under a Heaviside central HOD, no satellites, and a Press-Schechter mass function. Since the abstract advertises this argument as explaining the ELG versus LRG-HOD phenomenology, please validate the derivation against the actual HOD/HMQ weighting functions used in the mock catalogs, for example by direct numerical evaluation of Eq. (25) for the LRG-HOD-I, LRG-HOD-II, and HMQ parameter samples, to show that the simplifying assumptions do not drive the conclusion.","section":"Section 4.3, Eqs. (25)-(29)"}],"minor_comments":[{"comment":"The text states that 'the SFR cut selection provides a good approximation to samples of [OII] emitting galaxies' and refers to [62]; please add a brief caveat that this is an approximation and that the color-based selection is not used in the analysis, as the authors themselves note.","section":"Section 3.2"},{"comment":"The phrase 'sub Poisson stochasticity' should read 'sub-Poissonian stochasticity' for grammatical correctness.","section":"Section 4.1"},{"comment":"The sentence 'Astrid ELGs do not response monotonically to lowering of sSFR' contains a typo: 'response' should be 'respond'.","section":"Section 4.2"},{"comment":"When describing the HMQ normalizing-flow training, the paper says that αs and αc are excluded from the training sample; please clarify why these velocity bias parameters are excluded for ELG inference while they are included in the HMQ prior definition in Eq. (11).","section":"Section 4.4"},{"comment":"The captions should state explicitly that the LRG-HOD distributions are at z=0.5, the HMQ distributions at z=1.1, and the hydro galaxy measurements at z=0.5 (LRGs) and z=1 (ELGs), to avoid reader confusion about the redshift baselines.","section":"Figures 1 and 2"}],"recommendation":"major_revision","confidential_remarks":"The redshift mismatch in the central ELG versus LRG-HOD comparison and the Astrid kmax choice are concrete technical gaps that can likely be addressed with additional analysis. If the authors provide a same-redshift comparison or a quantitative redshift-evolution argument, and a robustness test for the Astrid scale cut, I would be supportive of publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a careful, useful validation paper, not a breakthrough. The genuinely new things are the field-level EFT parameters of DESI-like ELGs from MTNG and Astrid, the field-level comparison of hydro galaxies with HOD and HMQ mocks, and the normalizing-flow mapping from EFT to HOD parameters. The paper does most of the methodology right: it checks the cross-correlation coefficient, fits the noise model, uses the same field-level pipeline as the HOD prior work, and compares two hydro codes with different feedback physics. Credit where due: the analytic argument explaining why LRG b2 sits above the halo relation while ELG b2 sits on it is neat and, as far as I can see, correct in spirit. That is a useful result for anyone building EFT priors.\n\nThe main soft spot is real and worth naming: the central claim about ELG local bias being in tension with LRG-HOD models is tested at mismatched redshifts. The hydro ELGs are at z=1; the LRG-HOD predictive distributions are generated at z=0.5, and the HMQ mocks at z=1.1. Bias parameters evolve, so some of the \"tension\" could be a redshift effect. The paper does not generate LRG-HOD mocks at z=1 to close this gap. The analytic argument in Section 4.3 is suggestive but not a substitute. This does not kill the main validation of the HMQ model against the z=1 ELGs, but it weakens the strong claim that no LRG-like HOD can reproduce ELG bias. A same-epoch comparison would be straightforward to do with their existing pipeline.\n\nMinor things: Astrid redshift-space fits use kmax=0.4 rather than the 0.2 used elsewhere; the paper gives a reason (sample variance), but it makes the Astrid RSD comparison less apples-to-apples. The analytic b2 derivation is deliberately toy-like; fine, but it should not be oversold. No public release of catalogs or flows, only the Hi-Fi fitting code; for a validation paper, releasing the hydro galaxy catalogs would strengthen reproducibility.\n\nWho should read it: anyone building simulation-based priors for DESI full-shape analyses, and people calibrating HOD/HMQ models beyond the two-point function. I would send it to a serious referee with the request to check the same-epoch comparison before acceptance.","headline":"Solid field-level validation of HOD/HMQ against hydro galaxies, with a same-epoch gap in the ELG vs LRG-HOD comparison that should be fixed.","tokens_in":25571,"tokens_out":3275,"would_cite":true,"duration_ms":34246,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that red and blue galaxies in two large hydrodynamic simulations are accurately described by the field-level EFT model, with measured bias parameters matching the halo-occupation models used in survey analysis.","keywords":["effective field theory of large-scale structure","galaxy bias","halo occupation distribution","high mass quenched model","MillenniumTNG","Astrid simulation","redshift-space distortions","field-level forward model"],"falsifier":"Fit the full-shape power spectrum and bispectrum of DESI ELG data and measure b1 and b2: if b2 falls on the LRG-HOD relation rather than near the dark-matter halo curve, the paper's HMQ consistency claim and its HOD-weighting explanation for ELG bias would be wrong.","tokens_in":24484,"feed_emoji":"🔭","tokens_out":9022,"duration_ms":88922,"temperature":0.7,"pith_summary":"The paper tests whether the empirical halo-based recipes used to model galaxy clustering survive a harder test: comparing the full galaxy density field, not just the two-point function. Using the MillenniumTNG and Astrid hydrodynamic simulations, the authors select BOSS/DESI-like LRGs at z=0.5 and DESI-like ELGs at z=1, fit their fields with an EFT forward model, and compare the resulting bias and counterterm parameters with HOD and HMQ model distributions. They find that LRG parameters are consistent with LRG-HOD priors, while ELG local bias parameters are consistent with the HMQ model but not with LRG-like HOD models. They also find that ELGs show weaker fingers-of-God, which would make more of their data usable in perturbative analyses. If correct, this validates the halo-based priors used in full-shape cosmological analyses and extends them to the field level.","feed_headline":"Hydro galaxies match halo models at the field level","feed_subtitle":"Red and blue galaxy bias from two large simulations agrees with HOD/HMQ priors; ELGs show weaker fingers-of-God, promising more…","key_machinery":"The argument is carried by the field-level EFT forward model, which builds the galaxy density field from shifted bias operators (Zel'dovich-displaced linear density, tidal, and Galileon operators) and extracts transfer functions beta_i(k) directly from simulation snapshots, cancelling cosmic variance. These transfer functions are fitted with time-sliced perturbation theory to yield EFT bias parameters and redshift-space counterterms. The comparison side uses HOD and HMQ galaxy catalogs generated on the AbacusSummit small suite, processed with the same forward-model pipeline, and a normalizing flow trained on paired HOD-EFT samples converts measured EFT parameters back into inferred HOD parameter distributions.","core_discovery":"The paper establishes that the EFT-based field-level forward model accurately reproduces the large-scale density fields of hydrodynamic galaxies, and that the EFT parameters extracted from MTNG and Astrid match the predictions of the phenomenological halo-based models currently used for cosmological surveys. For red galaxies the match is to the decorated HOD models; for blue galaxies the match is to the HMQ model, which the paper shows reproduces both the anomalous local bias parameters and the weak fingers-of-God of ELGs. The authors also provide an analytic argument: the quadratic bias b2 of a galaxy sample is the HOD-weighted average of halo b2, so the shape of the HOD and the halo mass function determine whether galaxies lie above the halo b2(b1) curve, as LRGs do, or on it, as ELGs do.","pith_inferences":["An implication the authors leave implicit: if the weak ELG fingers-of-God are real, DESI's ELG sample could yield cosmological constraints comparable to or better than LRGs despite lower bias, because more modes enter the perturbative regime; this is directly testable with the DESI full-shape pipeline.","The b2(b1) distinction between LRG-like and ELG-like populations could be used as a data-driven diagnostic of the galaxy-halo connection: measuring b1 and b2 jointly in survey data tells you whether a target sample behaves like a thresholded high-mass HOD or a narrow-mass HOD without needing small-scale clustering.","A testable extension of the paper's logic is to build a suite of cosmological hydrodynamic simulations with varied subgrid feedback, which would quantify how much EFT parameter priors spread across plausible galaxy formation models and could then be marginalized over."],"forward_implications":["LRG-HOD priors used in EFT full-shape analyses of BOSS are robust to replacing halo mocks with hydrodynamic galaxies on quasi-linear scales.","The HMQ model can serve as a simulation-based prior for DESI ELG full-shape analyses, including higher-order bias parameters.","Because ELG fingers-of-God are weaker, the maximum wavenumber usable in ELG perturbative analyses can be larger, giving full-shape fits more usable modes.","The normalizing-flow mapping provides a cheap way to calibrate HOD parameters to reproduce the field-level clustering of hydrodynamic galaxies, and can be rerun for any new simulation or galaxy selection.","EFT parameter responses to selection cuts are not universal; they depend on the baryonic feedback model, so feedback must be controlled before using such priors blindly."],"supporting_citations":[{"why":"Provides the LRG-HOD-I sample, the field-level EFT fitting pipeline, and the HOD-based priors that the hydro LRG parameters are compared against.","marker":"[35]"},{"why":"Supplies the extended LRG-HOD-II priors with assembly and velocity bias and the normalizing-flow mapping between HOD and EFT parameters.","marker":"[36]"},{"why":"Introduces the field-level EFT forward model with shifted operators that the paper uses to extract transfer functions.","marker":"[43]"},{"why":"Establishes the transfer-function fitting procedure and the EFT error model used for the parameter measurements.","marker":"[49]"},{"why":"Provides previous EFT bias parameter measurements for IllustrisTNG galaxies that this work extends to MTNG and Astrid.","marker":"[57]"},{"why":"Defines the sSFR-based selection of BOSS/DESI-like samples in MTNG, the proxy on which the conclusions depend.","marker":"[62]"},{"why":"Provides the baseline ELG sSFR threshold selection used for the blue galaxy samples.","marker":"[64]"},{"why":"Defines the HMQ model for DESI ELGs that the blue galaxy EFT parameters are tested against.","marker":"[80]"},{"why":"Supplies the halo b2(b1) fit that anchors the analytic argument for why LRG b2 is enhanced while ELG b2 matches halos.","marker":"[84]"},{"why":"The AbacusSummit small suite on which the HOD and HMQ mock catalogs and their EFT parameter distributions are generated.","marker":"[107]"}],"fun_headline_variants":["EFT field model matches halo-based galaxy priors","ELGs show weaker fingers-of-God, EFT holds up","Simulation galaxies confirm halo-model bias predictions","Red and blue galaxy bias aligns with HOD and HMQ","Analytic b2 relation explains ELG vs LRG bias gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison treats galaxies selected by a specific star formation rate threshold at fixed number density as stand-ins for the observed BOSS LRGs and DESI ELGs; if those cuts do not faithfully match how surveys actually select these galaxies, the measured EFT parameters and the HOD/HMQ validation would not transfer to real data.","fun_headline_variants_meta":{"raw":{"variants":["EFT field model matches halo-based galaxy priors","ELGs show weaker fingers-of-God, EFT holds up","Simulation galaxies confirm halo-model bias predictions","Red and blue galaxy bias aligns with HOD and HMQ","Analytic b2 relation explains ELG vs LRG bias gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000305,"raw_usage":{"total_tokens":1791,"prompt_tokens":1025,"completion_tokens":766,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":641,"completion_tokens_details":{"reasoning_tokens":685}},"tokens_in":641,"tokens_out":766,"duration_ms":8804,"temperature":1.0,"reasoning_tokens":685,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:52:22.076855+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit the full-shape power spectrum and bispectrum of DESI ELG data and measure b1 and b2: if b2 falls on the LRG-HOD relation rather than near the dark-matter halo curve, the paper's HMQ consistency claim and its HOD-weighting explanation for ELG bias would be wrong.","supporting_citations":[],"review_version":1}