{"id":"31ea0767-4383-498b-9e41-bdc5a558095b","arxiv_id":"2501.14638","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Density-split correlation functions can be modeled analytically through the two-point density PDF; a large-deviation-theory version matches N-body simulations on large scales with a single fitted variance parameter.","lead":"Astrophysicists derived a formula for 'density-split' clustering, a way of measuring how galaxies cluster inside high- or low-density regions. The formula uses a large-deviation approximation and matches dark-matter simulations on large scales with only one fitted parameter, which could make cosmological measurements faster and more interpretable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The LDT bias is validated only against per-survey cosmic variance, not the 25-box mean error; a systematically wrong bias could pass while still being advertised as an analytical model.","rationale":"The reader's verdict is CONDITIONAL and already flags the lack of full independence in the validation, which overlaps with my concern. I agree that the large-separation assumption is a limitation, but the paper explicitly restricts its claim to s ≥ 40-50 Mpc/h, so that assumption is not a fatal flaw. The more load-bearing issue is the statistical power of the test: residuals are quoted in units of the per-survey cosmic variance, which is much larger than the error on the mean of the 25 simulation boxes. A model with a moderately wrong bias function can still satisfy the 'within cosmic variance of a DESI DR1 sample' criterion, so the validation does not establish that the LDT bias is correct to the precision implied by the 'analytical model' framing. This does not overturn the paper's stated claim, because the claim is explicitly about DESI-DR1-level accuracy, but it means the evidence is weaker than it first appears and the CONDITIONAL verdict remains appropriate. The proposed test is a concrete, computationally cheap way to settle whether the model is accurate at the simulation mean level; if it passes, the concern is resolved and the paper's central claim is strengthened.","tokens_in":23060,"tokens_out":16345,"duration_ms":159113,"concrete_test":"Recompute the residuals of the LDT density-split predictions in Figs. 6-8 dividing by the error on the mean, σ/√25 (or by the full covariance of the 25 boxes), instead of the per-box σ. If the model is discrepant at more than 2-3σ_mean in any density-split bin while remaining within 1σ_survey, the LDT bias has a systematic error hidden by the DESI-DR1-size tolerance; if it stays within 1-2σ_mean, the bias is validated at the simulation's own precision. The public code and the 25 mocks make this check straightforward.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the LDT bias factor predicts density-split correlations with one fitted parameter, σ_R. But the validation in Sec. 5.2.3 (Figs. 6-8) compares the model to the average of 25 AbacusSummit boxes, with residuals divided by the per-box standard deviation σ, i.e., the cosmic variance of a single DESI DR1-like survey. The error on the 25-box mean is σ/√25 ≈ σ/5. A residual of 1σ therefore corresponds to a roughly 5σ offset from the simulation mean. Since σ_R is fitted to the one-point PDF of these same boxes and ξ_R(s) is the ensemble-average smoothed correlation from the same boxes, the only genuinely predicted quantity is the separation-independent ratio B_DS = ∫_DS b^SN P^SN / ∫_DS P^SN. The test can thus pass even if the LDT bias shape is systematically wrong, provided the error is below the per-survey cosmic variance. This is a real soft spot for the claim that the LDT bias is theoretically correct and that the model is the first analytical description of density-split clustering: the evidence is consistent with the model being only an approximate fit within DESI DR1 noise, not with it being accurate at the simulation's own statistical precision.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops analytical expressions for density-split correlation functions, defined as the conditional expectation of the smoothed density field given that a spherical cell belongs to a density-split region. It derives the exact factorization of the density-split correlation into a bias factor times the smoothed two-point correlation function, and then presents three approximations for the required two-point density PDF: bivariate Gaussian, shifted log-normal, and a Large-Deviation-Theory (LDT) prediction based on spherical collapse. The models are validated against 25 AbacusSummit dark-matter boxes at z=0.8, including Poisson shot noise and a low-density sample mimicking the DESI DR1 LRG3+ELG1 number density. An extension to biased tracers is presented using Eulerian quadratic and Gaussian Lagrangian bias models with non-Poisson shot noise. The central claim is that the LDT model predicts the measured dark-matter density-split correlation functions on large scales (s≳40-50 h^-1 Mpc) within the cosmic variance of a DESI DR1-like survey, using only one fitted parameter, sigma_R.","tokens_in":23345,"tokens_out":7450,"duration_ms":93551,"significance":"If the LDT bias-factor prediction is validated at the statistical precision of the simulations themselves, the paper would provide a valuable analytic complement to simulation-based emulators for density-split clustering, and the explicit factorization in Eq. (4.7) is a useful formal contribution. The manuscript contains explicit derivations (Eqs. 3.7, 5.18, 5.35), public code, systematic comparisons between Gaussian, log-normal, and LDT models, and a nontrivial extension to biased tracers. The main limitations are that the LDT model's key amplitude input, the smoothed two-point correlation function, is measured from the same simulations used for validation, and that the statistical validation is performed against per-survey cosmic variance rather than the error on the 25-box mean. These issues do not invalidate the theoretical structure, but they currently leave the central claim less strongly supported than the abstract suggests.","major_comments":[{"comment":"The quantity called the LDT prediction is not self-contained: cξ_R(s) in Eq. (5.35) is the smoothed two-point correlation function measured from the same 25 AbacusSummit boxes, and σ_R is fitted to the one-point density PDF of those same boxes. Consequently, the only genuinely predicted part of Eq. (5.35) is the separation-independent bias-factor ratio in brackets; the overall amplitude and shape of ξ_DS are imported from the simulations. The statement in the abstract and §5.2.3 that the model relies on 'only one degree of freedom' is therefore misleading unless the measured ξ_R(s) input is explicitly counted. I recommend demonstrating the predictive power by using ξ_R(s) from a theoretical power-spectrum model or emulator (as is suggested for σ_R) or from an independent simulation suite, and/or by quantifying how sensitive the density-split predictions are to the choice of ξ_R(s). This is not circular in the strict sense, since the density-split measurements themselves are not used to fit parameters, but it is a data-input dependency that should be stated and tested.","section":"§5.2.3, Eq. (5.35), Fig. 6"},{"comment":"The validation panels divide residuals by the standard deviation over the 25 individual mocks, not by the error on the 25-box mean, which is σ/√25 ≈ σ/5. Because the model inputs (σ_R, cξ_R, and the one-point PDF) are themselves averages over the 25 boxes, the comparison against the average simulation measurement should use the mean error. Residuals of about 1σ in the bottom panels of Figs. 6-8 correspond to roughly 5σ offsets from the simulation mean. The statement that the model 'agrees with simulations on large scales within the cosmic variance of a typical DESI DR1 sample' is a statement about survey-level noise, but it does not establish that the LDT bias is accurate at the simulation's own statistical precision. Please add residual panels divided by σ/√25 and discuss the significance of any systematic residual trends.","section":"§5.2.3, Figs. 6-8 and Fig. 11"},{"comment":"The biased-tracer extension fits seven parameters (b1^E, b2^E, b1^G, b2^G, α0, α1, α2) to the same 25 ELG-populated AbacusSummit mocks that are then used to validate the model in Figs. 10 and 11. The agreement shown is therefore a fit to the validation data, not an independent prediction. The section is presented as an exploration, but if it is to support the claim of an analytical model for biased tracers in §6, a cross-validation procedure (e.g., fitting on a subset of boxes and predicting the rest, or fixing parameters from an independent HOD or galaxy sample) should be included. At minimum, the parameter counting should be stated explicitly in the text so readers can weigh the number of degrees of freedom against the quality of agreement.","section":"§6, Table 1 and Figs. 9-11"}],"minor_comments":[{"comment":"The bias estimator uses '3 out of the 6' positions separated by s; please specify exactly how the three directions are chosen and whether periodic boundary conditions are handled, for reproducibility.","section":"§4.4, Eq. (4.8)"},{"comment":"The text says δ0 is determined by the skewness relation but that in practice it is fitted; it would be useful to show explicitly how much the fitted value differs from the relation predicted by tree-order perturbation theory, since this is relevant to the comparison with LDT.","section":"§5.1.1, Eq. (5.5)"},{"comment":"The definition of σ_R,eff is implicit because P(ρ_R) depends on σ_R,eff through the exponential while σ_R,eff is defined by Eq. (5.30) with σ_R fitted externally. Please clarify the solution procedure (e.g., fixed-point iteration) and state explicitly that σ_R,eff is not an additional free parameter.","section":"§5.2.1, Eq. (5.30)"},{"comment":"For R=25 h^-1 Mpc the claimed validity range s≳50 h^-1 Mpc corresponds to s=2R, which is the non-overlap boundary, not the asymptotic s≫R regime; a sentence noting that the separation-independent bias approximation appears to work already at the non-overlap boundary would be helpful.","section":"§5.2.3, Fig. 8"},{"comment":"There are a few typographical issues (e.g., 'defintion' near Eq. (4.6) and inconsistent notation 'R1,R2' vs 'R_1,R_2'); a careful proofread would improve readability.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-organized and the theoretical derivations are largely explicit and checkable. My main reservation is that the strongest claim in the abstract ('agrees with simulations within the cosmic variance of a typical DESI DR1 sample despite relying on only one degree of freedom') is supported by a validation whose error bars are per-survey rather than per-simulation-mean, and whose amplitude input ξ_R(s) is measured from the same simulations. This is fixable within the scope of the paper by adding mean-error residual panels and by testing the model with ξ_R from an external or theoretical source, or at least by clearly qualifying the claim. The biased-tracer section should also be reframed as a fit unless a predictive test is added."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a serious look. This is the first analytical derivation of density-split correlation functions, connecting them to the two-point density PDF and bias function. The Gaussian, shifted log-normal, and LDT predictions are spelled out cleanly, and the LDT model outperforms log-normal with fewer free parameters. The paper is honest about its scope: real-space, top-hat smoothing, large separations, and the need for redshift-space and survey-specific extensions.\n\nWhat is actually new: the factorization of the density-split correlation into a bias factor times the smoothed two-point correlation, and the derivation of that bias factor under three approximations. The LDT bias function itself comes from earlier work by Codis et al., but the connection to density-split clustering is new and the presentation is transparent. The code is public, which helps reproducibility.\n\nSoft spots, in proportion. First, the validation is not fully blind: σ_R is fitted to the one-point PDF of the same AbacusSummit boxes, and ξ_R(s) is measured from the same boxes. The genuinely predicted quantity is the ratio B_DS, and that ratio is the only place the LDT model could fail. Second, the stress-test concern is legitimate: the residual plots divide by per-survey cosmic variance (the standard deviation of the 25 boxes, not the error on the mean). A systematic offset smaller than that per-survey variance would pass the test. The paper's claim is phrased carefully as \"within the cosmic variance of a typical DESI DR1 sample,\" which is the right target for survey application, but it does not establish that the LDT bias is accurate at the simulation's own statistical precision. For a claim of theoretical correctness, that distinction matters. It is a soft spot, not a fatal one.\n\nThe biased-tracer section fits seven parameters and is more of an exploratory extension than a sharp test; the Gaussian Lagrangian bias works reasonably, but the fits are not the core of the paper. The Gaussian approximation fails as expected, which is fine as a pedagogical baseline.\n\nWho is this for? Anyone working on density-split clustering or counts-in-cells statistics, especially for DESI, Euclid, or LSST forecasting and analysis. It gives a fast, interpretable alternative to emulators on large scales, and it identifies where the information resides. The limitations are stated in the paper, not hidden.\n\nI would send this to peer review. It is a solid, reproducible contribution with a clear caveat about validation precision that a referee should ask the authors to address. A natural request would be to show residuals against the 25-box mean error, or to use a subset of boxes for fitting σ_R and a separate subset for validation.","headline":"A genuinely useful analytical model for density-split clustering, with an honest but not fully closed validation: the LDT bias matches simulations at DESI-DR1-like precision, but the test does not establish accuracy at the simulation's own statistical precision.","tokens_in":23937,"tokens_out":1141,"would_cite":true,"duration_ms":12931,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Density-split galaxy clustering can now be modeled analytically, with one fitted parameter, at separations above about 40 Mpc/h.","keywords":["density-split clustering","large deviation theory","count-in-cell statistics","two-point correlation function","galaxy bias","cosmological large-scale structure","N-body simulations","DESI"],"falsifier":"Take one N-body simulation at $R=10\\,h^{-1}\\mathrm{Mpc}$, measure the density-split correlation at $s=50\\,h^{-1}\\mathrm{Mpc}$, and recompute the LDT prediction using $\\sigma_R$ from a non-linear power-spectrum emulator instead of fitting it; the paper's claim that one (or zero) parameter suffices predicts agreement within cosmic variance, so a clear offset would falsify the central claim.","tokens_in":2060,"feed_emoji":"🌌","tokens_out":2385,"duration_ms":75049,"temperature":0.7,"pith_summary":"The paper sets out to replace simulation-based emulators for density-split clustering, a statistic that measures how galaxies cluster in different density environments, with a first-principles analytical model. It shows that the density-split correlation function factorizes into a bias factor times the smoothed two-point correlation function, so all extra information beyond the standard two-point statistic sits in the bias factor. The paper derives this bias factor under three approximations (Gaussian, shifted log-normal, and large-deviation theory) and finds that the large-deviation version, built from spherical-collapse dynamics, matches dark-matter N-body simulations within cosmic variance at separations $s \\gtrsim 40\\text{--}50\\,h^{-1}\\mathrm{Mpc}$ using only one fitted parameter, the smoothed density variance $\\sigma_R$. If correct, this gives a fast, physically interpretable model for large-scale density-split clustering in surveys such as DESI.","feed_headline":"One parameter predicts density-split clustering on large scales","feed_subtitle":"Large-deviation theory replaces simulation emulators for density-split statistics above 40 Mpc/h.","key_machinery":"The load-bearing object is the bias function and its factorization. In the large-separation limit the two-point density PDF is written as $P(\\delta_R,\\delta'_R)=P(\\delta_R)P(\\delta'_R)[1+\\xi_R(s)\\,b(\\delta_R,s)\\,b(\\delta'_R,s)]$, and the density-split correlation function becomes $\\xi^{\\mathrm{DS}}_R(s)=B_{\\mathrm{DS}}\\,\\xi_R(s)$ with $B_{\\mathrm{DS}}=\\int_{\\mathrm{DS}} b(\\delta_R)\\,P(\\delta_R)\\,d\\delta_R\\,/\\,\\int_{\\mathrm{DS}} P(\\delta_R)\\,d\\delta_R$. The LDT input is the bias function $b(\\rho_R)=\\tau(\\rho_R)/\\sigma^2_{r,L}$, where $\\tau$ is the initial density contrast mapped to the evolved density $\\rho_R$ by the spherical-collapse relation $\\rho_R\\simeq(1-\\tau/\\nu)^{-\\nu}$ with $\\nu=21/13$; $\\sigma_{r,L}$ is the variance of the initial field at the Lagrangian radius $r=R\\rho_R^{1/3}$. This gives a separation-independent bias in the regime $s\\gg R$, and Poisson shot noise is included by convolving the PDFs. Nearly all physical content is carried by this bias function, which is why the model has only one fitted parameter, $\\sigma_R$.","core_discovery":"The central claim is that density-split correlation functions can be expressed as a bias factor multiplying the smoothed two-point correlation function $\\xi_R(s)$, with the bias factor built from the one-point density PDF and the density bias function. In the large-separation limit the large-deviation-theory (LDT) bias function $b(\\rho_R)=\\tau(\\rho_R)/\\sigma^2_{r,L}$ becomes independent of separation, and when inserted into the factorization it predicts the measured density-split correlation functions of N-body simulations within cosmic variance for $s \\gtrsim 40\\text{--}50\\,h^{-1}\\mathrm{Mpc}$ with one free parameter, $\\sigma_R$. The same LDT machinery is extended to biased tracers by combining the matter PDF with a Gaussian Lagrangian bias and non-Poisson shot noise, matching ELG-populated simulations on large scales. The paper also demonstrates that the Gaussian approximation fails and that the shifted log-normal model, despite fitting the one-point PDF well, does not reach the accuracy of LDT for the density-split correlation functions.","pith_inferences":["Beyond the paper: if the factorization and LDT bias hold for real galaxies, density-split clustering becomes a cheap, interpretable probe for DESI-like surveys; the main obstacle is redshift-space distortions, which the paper leaves to future work.","Beyond the paper: the scale at which the bias becomes separation-independent could itself be measured and used as a consistency test of spherical-collapse dynamics.","Beyond the paper: the LDT machinery may naturally extend to beyond-$\\Lambda$CDM physics such as neutrino mass, modified gravity, or primordial non-Gaussianity, by replacing the spherical-collapse mapping with modified collapse dynamics and comparing to simulations of those models."],"forward_implications":["On large scales ($s \\gtrsim 40\\text{--}50\\,h^{-1}\\mathrm{Mpc}$), density-split correlation functions can be predicted without a simulation-based emulator, using one fitted parameter $\\sigma_R$ or, if $\\sigma_R$ comes from a power-spectrum emulator, no fitted parameters at all.","All information in density-split clustering beyond the standard two-point correlation function is contained in the scale-independent bias factor, providing a physical explanation for what density-split statistics add.","The shifted log-normal approximation, although excellent for the one-point PDF, is not sufficient for density-split correlations at the statistical precision of current mocks; LDT is required.","The model extends to biased tracers: combining the LDT matter PDF with a Gaussian Lagrangian bias and non-Poisson shot noise matches ELG-populated N-body simulations within cosmic variance for $s \\gtrsim 40\\,h^{-1}\\mathrm{Mpc}$.","Because the bias factor is separation-independent in the large-scale regime, the model can serve as a fast theoretical ingredient for survey analyses and as an initial guess to improve emulator-based approaches."],"supporting_citations":[{"why":"Supplies the large-separation LDT bias function $b(\\rho_R)=\\tau(\\rho_R)/\\sigma^2_{r,L}$ that the model uses as its central ingredient.","marker":"[34]"},{"why":"Provides the LDT one-point PDF for the log-density in spheres, extended to variances near 1, which the model adopts and normalizes.","marker":"[31]"},{"why":"Establishes the bias-function description of the two-point density PDF and the normalization constraints $\\langle b\\rangle=0$ and $\\langle\\rho_R b\\rangle=1$.","marker":"[32]"},{"why":"Provides the LDT rate-function implementation and the low-density bias prediction that the paper adapts for the density-split model.","marker":"[33]"},{"why":"Supplies the shifted log-normal joint model and Poisson shot-noise convolution used as a competing approximation.","marker":"[27]"},{"why":"Derives the log-normal bias function in the large-separation limit, used as the log-normal model's bias.","marker":"[43]"},{"why":"The simulation-based density-split model that this paper's analytical model aims to replace on large scales.","marker":"[23]"},{"why":"Provides the non-Poisson shot-noise model and Gaussian or Eulerian bias fits used in the biased-tracer extension.","marker":"[50]"}],"fun_headline_variants":["Large-deviation theory predicts density-split clustering with one parameter","One parameter explains density-split clustering on large scales","LDT outperforms log-normal for density-split clustering","Analytical model for density-split clustering from one parameter"],"cache_read_input_tokens":25984,"weakest_assumption_plain":"The prediction rests on the assumption that at separations large compared with the smoothing radius the bias factor is exactly the large-separation LDT bias $b(\\rho_R)=\\tau(\\rho_R)/\\sigma^2_{r,L}$, meaning that spherical collapse gives the most likely mapping from initial to evolved density and that the bias is independent of separation; the paper only validates this for $s\\gtrsim40\\text{--}50\\,h^{-1}\\mathrm{Mpc}$ and does not test overlapping spheres ($s<2R$).","fun_headline_variants_meta":{"raw":{"variants":["Large-deviation theory predicts density-split clustering with one parameter","One parameter explains density-split clustering on large scales","LDT outperforms log-normal for density-split clustering","Analytical model for density-split clustering from one parameter"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000407,"raw_usage":{"total_tokens":2111,"prompt_tokens":942,"completion_tokens":1169,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":1102}},"tokens_in":558,"tokens_out":1169,"duration_ms":7921,"temperature":1.0,"reasoning_tokens":1102,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:57:10.321404+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one N-body simulation at $R=10\\,h^{-1}\\mathrm{Mpc}$, measure the density-split correlation at $s=50\\,h^{-1}\\mathrm{Mpc}$, and recompute the LDT prediction using $\\sigma_R$ from a non-linear power-spectrum emulator instead of fitting it; the paper's claim that one (or zero) parameter suffices predicts agreement within cosmic variance, so a clear offset would falsify the central claim.","supporting_citations":[{"cited_title":"Codis, F","cited_arxiv_id":null,"evidence_quote":"Supplies the large-separation LDT bias function $b(\\rho_R)=\\tau(\\rho_R)/\\sigma^2_{r,L}$ that the model uses as its central ingredient."},{"cited_title":"Uhlemann, S","cited_arxiv_id":null,"evidence_quote":"Provides the LDT one-point PDF for the log-density in spheres, extended to variances near 1, which the model adopts and normalizes."},{"cited_title":"Uhlemann, S","cited_arxiv_id":null,"evidence_quote":"Establishes the bias-function description of the two-point density PDF and the normalization constraints $\\langle b\\rangle=0$ and $\\langle\\rho_R b\\rangle=1$."},{"cited_title":"Codis, C","cited_arxiv_id":null,"evidence_quote":"Provides the LDT rate-function implementation and the low-density bias prediction that the paper adapts for the density-split model."},{"cited_title":"Friedrich, D","cited_arxiv_id":null,"evidence_quote":"Supplies the shifted log-normal joint model and Poisson shot-noise convolution used as a competing approximation."},{"cited_title":"Uhlemann, O","cited_arxiv_id":null,"evidence_quote":"Derives the log-normal bias function in the large-separation limit, used as the log-normal model's bias."},{"cited_title":"Cuesta-Lazaro, E","cited_arxiv_id":null,"evidence_quote":"The simulation-based density-split model that this paper's analytical model aims to replace on large scales."}],"review_version":1}