{"id":"282d6ac5-575e-4154-8d62-f065cd4aa825","arxiv_id":"2601.05934","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"In f(R) simulations, environment-split marked correlation functions show the strongest modified-gravity signal in filaments and nodes, but the claimed factor-of-four information gain depends on an undefined and statistically questionable chi-square normalization.","lead":"This paper splits simulated galaxies by cosmic web environment (nodes, filaments, walls, voids) and asks whether marked clustering statistics become more sensitive to f(R) modified gravity. It reports that filaments and nodes add the most information, claiming up to a factor-of-four gain in constraining power for CMASS- and DESI-like mocks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Information-gain claims rest on an undefined χ2_ν; Figure 7's factors cannot be audited, and the all-galaxies baseline may not be computed with the same pipeline.","rationale":"The reader correctly identifies that the χ2_ν statistic is never defined and that the headline information-gain factors depend entirely on it. This is a load-bearing concern: the abstract and conclusions quantify the improvement as a factor of two and four, and Figure 7 is the only evidence. However, the reader's specific mathematical argument that a concatenated reduced chi-square cannot exceed the all-galaxies value is only valid under a block-diagonal covariance. If the combined data vector includes cross-environment covariance terms, the reported factors can in principle be genuine, even when each component individually has a lower reduced chi-square than the all-galaxies baseline. The paper does not provide the formula for the combined covariance or the cross-covariance blocks, so the reader's strong 'cannot' overstates the case. Nevertheless, the absence of an explicit definition and degrees-of-freedom count means the claim cannot be audited. Additionally, the all-galaxies baseline is described as coming from Armijo et al. (2018), raising a potential pipeline mismatch. These issues together undermine the central quantitative claim, supporting a reject verdict: the paper as written does not establish the stated information gains. A revision providing the definition, covariance construction, and a recomputed baseline could change the verdict, but the current version is not reproducible.","tokens_in":16001,"tokens_out":12602,"duration_ms":140456,"concrete_test":"Ask the authors to provide the exact definition of χ2_ν, the number of data points per vector, and the covariance matrix construction for combined data vectors (Eq. A1 blocks). Independently recompute Figure 7 from the supplied data vectors: compute δ = M_F5 − M_GR and χ2_ν = δ^T C^{-1} δ / ν for each environment, the all-galaxies vector, the concatenated vector, and the node+filament vector. Check whether the reported factors 2.3 and 4.1 are reproduced when the all-galaxies baseline is recomputed with the same jackknife procedure (not taken from Armijo et al. 2018). If the factors vanish or change materially after this audit, the information-content claim is unsupported.","verdict_should_be":"REJECT","load_bearing_attack":"In §5.3, after defining (S/N)^2 in Eq. 16, the authors state they calculate 'the individual reduced chi-square (χ2_ν)' but never give its formula or the number of degrees of freedom. Figure 7 then reports χ2_ν for individual environments, all-galaxies, and combined vectors, and the abstract/conclusions quote factors of 2.3 and 4.1. This is the sole quantitative support for the headline claim. The reader's objection that a concatenated reduced chi-square is a weighted average and hence cannot exceed all components is only true if the covariance is block-diagonal; with cross-environment covariance the combined value can exceed each component. But the paper neither states how the combined covariance is built nor shows the cross-blocks, so we cannot tell whether the factors reflect physical information or normalization. Moreover, the text says the all-galaxies benchmark comes 'from (Armijo et al. 2018)' — if those values were not recomputed with the same jackknife covariance, binning, and mark pipeline, the comparison is uncontrolled. A standard reduced chi-square for a detection (δ = F5−GR) would be χ2/ν; if ν is the number of bins, a combined vector with four times more bins trivially increases raw χ2, so the factor-of-two claim depends critically on the normalization. Without the definition and the actual covariance matrices, the central claim is not reproducible.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses two large N-body simulations, GR and Hu-Sawicki f(R) gravity with |f_R0|=10^-5, classifies the cosmic web into nodes, filaments, walls, and voids with pycosmommf, populates haloes with HOD parameters matched to CMASS- and DESI-like number densities and clustering, and computes density- and mass-marked correlation functions in each environment at z=0 and z=1. The main claim is that environmental splitting, especially filaments and nodes, enhances sensitivity to f(R) gravity, and that combining environmental data vectors raises the reduced chi-square by a factor of 2.3 relative to the all-galaxy measurement, with node+filament giving a factor of 4.1. These factors, taken from Fig. 7, are the quantitative basis for the abstract and conclusions.","tokens_in":16313,"tokens_out":5266,"duration_ms":59165,"significance":"If substantiated, the result would be practically important: it suggests that a cheap post-processing step—splitting galaxy samples by cosmic web environment—can roughly double the constraining power of marked correlation functions for modified gravity, and that filaments are not just the largest-volume environment but also carry the most additional information. The paper has clear strengths: it uses well-defined simulation pairs, matches HOD mocks to the two-point clustering, provides jackknife covariance information in Appendix A, and the qualitative differences between GR and F5 in Figures 4 and 5 are visually credible. However, the central quantitative claim about factor 2.3 and 4.1 improvements rests entirely on an undefined reduced chi-square, so the headline numbers are not currently auditable. The paper would be significantly stronger if the metric were precisely defined and the all-galaxy baseline recomputed in the same pipeline.","major_comments":[{"comment":"The quantity χ2ν is never defined. Eq. (16) defines (S/N)^2, but the text then states that 'the individual reduced chi-square (χ2ν)' is calculated without giving the data vector, the number of degrees of freedom, or the covariance matrix used for individual versus combined vectors. The abstract's headline factors of 2.3 and 4.1 come entirely from Fig. 7. A reduced chi-square for a concatenated vector can exceed the component reduced chi-squares when cross-environment covariance is included, so the quoted behavior is not by itself impossible; however, without the formula and the combined covariance blocks, the reader cannot tell whether the factors reflect physical information or simply the larger number of bins/parameters in the combined vector. Please state the exact formula for χ2ν, report the degrees of freedom, and show the covariance matrix for each combined vector; also give the eq","section":"§5.3, Eq. (16), Fig. 7"},{"comment":"The baseline values in Fig. 7 are quoted as 'results of (Armijo et al. 2018)' rather than recomputed with the same mocks, binning, mark definitions, and jackknife covariance used for the environment-split measurements. Since this baseline appears in the denominator of the headline 'factor of two / factor of four' claims, any difference in pipeline can create an apparent gain. The all-galaxy χ2ν must be recomputed with the identical pipeline and compared on equal footing.","section":"§5.3, all-galaxy benchmark"},{"comment":"The interpretation of χ2ν as 'information content' is problematic. A reduced chi-square is a goodness-of-fit statistic, not a detection significance or Fisher information. The values in Fig. 7 are hard to interpret without degrees of freedom: χ2ν=800 for the DESI-like node+filament combination could be an extremely strong detection if ν is small, or a poor fit if ν is large. The paper should report a detection statistic with fixed degrees of freedom (e.g., Δχ2 between F5 and GR, or the S/N from Eq. 16) and state explicitly whether the quoted factors are in Δχ2, χ2ν, or S/N.","section":"§5.3, Fig. 7, conclusions"}],"minor_comments":[{"comment":"The cell size used for the density field is L_cell=2.19 Mpc/h in §3.2 but L_cell=2 Mpc/h in §4. Please use one consistent value and explain the difference if intentional.","section":"§3.2 and §4"},{"comment":"The sentence 'whereas the can also have an impact in the error bars' is incomplete; it should refer to the random seed uncertainty.","section":"§3.4"},{"comment":"The mock samples are at z=0 and z=1 while CMASS and DESI LRG samples have effective redshifts of about 0.5 and 0.8. This is acknowledged, but the forecasts should be labeled as approximate and the mismatch discussed quantitatively.","section":"§3.4"},{"comment":"Several placeholder citations appear as '(cite)' or '(cites)' (e.g., §3.3, §3.4). These need to be resolved.","section":"Throughout"},{"comment":"The appendix shows correlation coefficients for HOD2 only; for reproducibility, include the actual covariance matrices (or a link to them) for all samples and for the combined vectors used in Fig. 7.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"I do not see a novelty or misconduct issue. The central concern is statistical: the headline information-gain factors rest on an undefined χ2ν and an uncontrolled all-galaxy baseline. This is fixable within the scope of the paper by recomputing or clearly defining the statistic and its covariance, so I recommend major revision rather than rejection. If the reanalysis shows the factors become much smaller, the conclusions should be revised accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper has a genuinely useful qualitative result buried under a quantitative claim that cannot be audited. The environment-split marked correlation functions do show a consistent F5/GR pattern — nodes and filaments deviate more than walls and voids — and the authors did real work building HOD mocks matched in number density and clustering, using the Arnold et al. f(R) simulations and pycosmommf classification. If the question is whether cosmic-web environment changes the response of marked statistics to f(R) gravity, the figures answer yes, and I see no reason to doubt that. That part deserves credit.\n\nThe soft spot is the headline. Section 5.3 says they compute the individual reduced chi-square (chi2_nu) but never gives the formula, the number of degrees of freedom, or how the covariance of a combined data vector is assembled. Figure 7 reports values like 400 and 800 for the DESI-like sample. A reduced chi-square with maybe a few dozen bins cannot be 400 unless the model is catastrophically wrong; what is plotted looks like raw chi2 or (S/N)^2 without per-dof normalisation. If that is the case, the factors 2.3 and 4.1 quoted in the abstract mostly reflect the fact that concatenating four data vectors adds up raw sums, not that there is four times more physical information. The baseline is also uncontrolled: the all-galaxies curve is cited to Armijo et al. 2018, but the paper does not say it recomputed that benchmark with the same jackknife covariance, binning, mark definition, and HOD mocks. The stress-test note is right that a combined reduced chi-square can exceed every component when cross-environment covariance is non-zero, so the reader's simplistic weighted-average objection is not decisive on its own. That cuts both ways: because the covariance blocks are not shown or described, we cannot tell whether the improvement is real or a normalization artifact. This is the central claim, so it has to be fixed before the paper is usable.\n\nSmaller issues: F5 curves in Figures 4 and 5 have no error bars, so visual deviations up to 20% are not quantified; and the abstract and conclusions compare against slightly different baselines. Both are secondary.\n\nBottom line: this is a sensible, competent measurement paper that currently overstates what it demonstrates. I would send it to a referee rather than desk reject, because the idea is timely and the qualitative signal is credible. The referee should require an explicit definition of chi2_nu, the dof, the jackknife covariance construction for combined vectors, and a same-pipeline all-galaxies baseline. Until then I would not cite the information-gain numbers.","headline":"Environmental splitting of marked statistics is a sensible idea with credible qualitative plots, but the headline 'factor of 2.3/4.1 information gain' rests on an undefined chi2_nu and an uncontrolled baseline.","tokens_in":16823,"tokens_out":3607,"would_cite":false,"duration_ms":39911,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that weighting galaxy pairs by cosmic-web environment—especially nodes and filaments—roughly doubles, and up to quadruples, the sensitivity of marked correlation functions to f(R) modified gravity relative to measuring the","keywords":["modified gravity","f(R) gravity","marked correlation function","cosmic web","galaxy clustering","large-scale structure","cosmic filaments","chameleon screening"],"falsifier":"Take the quoted per-environment reduced chi-square values and the quoted combined values, and check whether the combined value is a weighted average of the components divided by the total number of bins. If the reduced chi-square is defined in the standard way, a concatenated data vector cannot exceed every component; finding that it does would show the gain is an artifact of bin counting rather than information content.","tokens_in":15852,"feed_emoji":"🌌","tokens_out":7698,"duration_ms":84977,"temperature":0.7,"pith_summary":"This paper tries to establish that the cosmic web provides a new lever arm for testing modified gravity. Using marked correlation functions—pair-counting statistics in which each galaxy is weighted by local density or host halo mass—the authors split a large f(R)-gravity simulation into nodes, filaments, walls, and voids, and compare with general relativity. They find that the f(R) signal is strongest for galaxies in nodes and filaments, and that combining environment-split measurements raises the discriminating power by about a factor of two over the standard all-galaxy marked correlation function, with the node-plus-filament combination exceeding it by a factor of four. If the claim holds, survey analyses need not rely only on overall clustering; splitting by environment is a cheap way to sharpen constraints on deviations from general relativity.","feed_headline":"Splitting galaxies by cosmic web doubles f(R) gravity signal","feed_subtitle":"Weighting pairs by cosmic environment, especially filaments, sharpens tests of modified gravity.","key_machinery":"The workhorse is the marked correlation function, M(r) = (1 + W(r))/(1 + ξ(r)), where ξ is the standard two-point correlation function and W is the same pair count weighted by a mark per galaxy. The marks used are local density contrast raised to the power p = 0.5 and host halo mass raised to the power p = 0.5. Environment labels are assigned by sorting the eigenvalues of the Hessian of a smoothed density field into nodes (all positive), filaments (two positive), walls (one positive), and voids (all negative), thereby selecting the regions where the f(R) fifth force is unscreened. The ratio form matters because it suppresses galaxy-bias and selection effects while retaining the environmental","core_discovery":"On its own terms, the paper claims that the environmental structure of the cosmic web is not just a source of nuisance but a carrier of signal for modified gravity. In f(R) gravity with a present-day scalaron amplitude of |f_R0| = 10^-5, the fifth force is screened in high-density nodes but active in lower-density environments, and the simulations show that this leaves measurable imprints in the marked correlation function: deviations from GR of up to about 20 percent for galaxies in nodes and filaments on scales of roughly 1 to 20 Mpc/h. The decisive quantitative claim is that measuring the marked correlation function separately for nodes, filaments, walls, and voids, then concatenating the","pith_inferences":["The quoted factor-of-2.3 and 4.1 comparisons rely on the unreported definition of the reduced chi-square statistic; until the degrees-of-freedom accounting is specified, the gain should be read as a comparison of the plotted summary values rather than a proven information-theoretic result.","The same environment-splitting strategy could be exported to other screening mechanisms and to other statistics, such as marked power spectra or density-split clustering; a filament-specific gain there would corroborate that the web environment, rather than the choice of mark, is the active ingredient.","A sharper test would recompute the environment-split marked correlation function using covariance from many independent simulation realizations rather than jackknife subvolumes; if the factor-of-two gain survives that covariance treatment, the forecast is on firmer ground."],"forward_implications":["Environment-split marked correlation functions raise the expected signal-to-noise for low- and high-redshift LRG-like mocks by about a factor of two over the all-galaxy marked correlation function; combining nodes and filaments alone exceeds it by a factor of four.","Filaments, not just voids, are a primary carrier of f(R) information at non-linear scales (roughly 1 to 5 Mpc/h), while voids contribute mainly at large scales beyond about 20 Mpc/h.","Both density-based and host-halo-mass-based marks separate the F5 model from GR, with relative residuals up to about 20 percent for nodes and filaments at small separations.","The constraining power in a higher-redshift, higher-number-density LRG-like sample is about four times larger than in a lower-redshift, lower-density sample.","The environment classification is meaningful in both GR and f(R) simulations, and galaxies in unscreened environments are systematically more massive in the f(R) model."],"fun_headline_variants":["Cosmic web environments double f(R) gravity signal","Filaments and nodes sharpen modified gravity tests","Marked pairs in cosmic web boost f(R) signal twofold","Environmental marks reveal f(R) gravity in clustering","Cosmic web weighting doubles sensitivity to f(R) gravity"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire information-gain claim hangs on how the quoted chi-square values are defined and normalized, and the paper never gives that definition; under the standard definition of reduced chi-square, the reported factors of 2.3 and 4.1 would be hard to reconcile with the smaller per-environment values.","fun_headline_variants_meta":{"raw":{"variants":["Cosmic web environments double f(R) gravity signal","Filaments and nodes sharpen modified gravity tests","Marked pairs in cosmic web boost f(R) signal twofold","Environmental marks reveal f(R) gravity in clustering","Cosmic web weighting doubles sensitivity to f(R) gravity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000554,"raw_usage":{"total_tokens":2448,"prompt_tokens":691,"completion_tokens":1757,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":1681}},"tokens_in":435,"tokens_out":1757,"duration_ms":12731,"temperature":1.0,"reasoning_tokens":1681,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T11:29:44.109474+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the quoted per-environment reduced chi-square values and the quoted combined values, and check whether the combined value is a weighted average of the components divided by the total number of bins. If the reduced chi-square is defined in the standard way, a concatenated data vector cannot exceed every component; finding that it does would show the gain is an artifact of bin counting rather than information content.","supporting_citations":[],"review_version":1}