{"id":"4a9116a6-e411-4860-8e97-d2c11bcc17b3","arxiv_id":"2502.09709","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Betti curves and kNN distributions give similar cosmological constraints in Quijote simulations, with beta0/beta1 dominating Betti information and the two statistics only partially redundant.","lead":"This paper compares how much cosmological information Betti curves and k-nearest neighbor distributions extract from simulated halo catalogs. It finds they are roughly competitive on non-linear scales and not fully independent, which matters for choosing summary statistics in future galaxy surveys.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The β_d > 5×10^-4 tail cut (Sec. 5.2) may remove the very β2 bins that carry its signal, so the claim that β2 has almost no constraining power is not yet robust.","rationale":"The central claim is a ranking of information content, and Eq. (3) is the tool that produces that ranking, so the Gaussianity plus fixed-covariance assumption is load-bearing. The paper's safeguards are the tail cut and a qualitative χ² test; both are too weak to guarantee the assumption in the low-amplitude β2 tail. I agree with the reader's weakest_assumption but sharpen it to a concrete mechanism: the 5×10^-4 threshold is applied per Betti curve, and for β2 this threshold is not a far tail—it removes a substantial fraction of the curve's large-radius bins. Because the abstract's strongest claim (β2 ≈ no constraining power) concerns exactly this curve, the cut can decide the conclusion. A secondary issue is that the same statement is parameter-dependent: in the {Ω_m, σ8, n_s} space β2 does constrain n_s (Sec. 7), so the abstract's wording overgeneralizes. The body is careful to limit the main claim to the 2D space, which is why I do not recommend changing the reader's conditional verdict; instead, the conditionality should be made explicit and the tail-cut robustness demonstrated.","tokens_in":15247,"tokens_out":9793,"duration_ms":100714,"concrete_test":"Recompute the β2-only Fisher matrix without the β_d > 5×10^-4 cut (or with the threshold lowered by a factor of 10) and with a Gaussianized data vector (e.g., a per-bin rank or Box-Cox transform), using the same Quijote derivatives and covariance. If the β2 figure-of-merit relative to β0+β1 changes by more than about 50%, the central claim that β2 has almost no constraining power is not robust to the Gaussianity assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The entire comparison rests on Eq. (3), which assumes a multivariate Gaussian likelihood with a parameter-independent covariance. To enforce this, Section 5.2 discards all Betti bins with β_d ≤ 5×10^-4, and Appendix A offers only a qualitative χ² test against a Gaussian null. For β2, the curve decays toward zero at large r, so the 5×10^-4 cut removes precisely the large-scale tail of the curve; the surviving β2 bins are small-amplitude and likely noise-dominated. The χ² test with roughly 1,500 simulations and 60–150 degrees of freedom has limited power to detect non-Gaussianity in these low-count bins, so 'almost no constraining power comes from β2' (abstract) may be an artifact of the cut rather than a property of β2. The same issue applies to the kNNs: Section 7 says tails were removed, but Section 5.3 specifies no cut, making the comparison protocol ambiguous. If the discarded tails carry non-Gaussian cosmological information, the FoM ranking in Table 1 and the conclusion that kNNs and Betti curves are not fully independent could change. This concern is compounded by the paper's own finding that β2 does constrain n_s in the 3-parameter space (Sec. 7, Fig. 9), so the blanket statement that β2 is uninformative is parameter-dependent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares the cosmological information content of Betti curves and k-nearest neighbor (kNN) distributions as summary statistics for halo clustering, using Fisher matrices estimated from the Quijote simulations on scales 2–50 h^-1 Mpc. The authors pay careful attention to convergence: they use 5,000 fiducial simulations for covariance, apply the Hartlap correction, and diagnose derivative-noise convergence through an eigenvalue decomposition, concluding that only two parameter directions are reliably constrained. Restricting to {Omega_m, sigma8}, they find that beta0 and beta1 contain nearly all the Betti-curve information while beta2 contributes little, that DD-kNNs are highly competitive while DR-kNNs are less so, and that the kNN and Betti probes are complementary but not fully independent.","tokens_in":15553,"tokens_out":10104,"duration_ms":102144,"significance":"If the main results hold, the paper provides practical guidance for survey analysis: Betti-curve analyses could focus on beta0/beta1, and kNNs could serve as a cheaper substitute or complement. The Fisher convergence methodology is a genuine strength: the use of 5,000 simulations, Hartlap correction, eigenvalue-based noise fractions, and the explicit restriction to converged parameter directions is careful and reproducible with public Quijote data. The work also highlights a possible connection between topological statistics and kNNs. However, the headline claims rest on several choices—the Betti tail cut, the unspecified treatment of kNN tails, and the way probe combinations are constructed—that are not yet fully validated, so the quantitative conclusions should be treated with caution until those choices are tested.","major_comments":[{"comment":"The claim that beta2 has almost no constraining power is not robust to the ad hoc tail cut beta_d > 5e-4. Because beta2 decays toward zero at large r (Fig. 2), this cut removes exactly the large-scale tail of the beta2 curve; the surviving bins are low-amplitude and potentially noise-dominated. The Gaussianity check in Appendix A is qualitative (a chi-square histogram with roughly 1,500 simulations and 60-150 degrees of freedom), so it has limited power to detect non-Gaussianity in these low-count bins. The paper should show that the beta2 conclusion is stable to the threshold (for example, by repeating the Fisher calculation for several values of the cut) or quantify the information contained in the discarded tail; as written, the abstract claim that 'almost no constraining power comes from beta2' may be an artifact of the cut. The parameter dependence is also visible in §7 and Fig. 9, where beta2 does constrain n_s, so the statement should be qualified.","section":"§5.2, Fig. 2, App. A"},{"comment":"The method used to build Fisher matrices for probe combinations is not stated. If each combination is computed as the sum of the individual Fisher matrices, then independence is assumed by construction and the conclusion that the probes are 'not fully independent' is circular. To test independence, the authors need to compute the joint Fisher matrix from the concatenated data vector (including the cross-covariance between probes) and compare it with the sum of the individual Fisher matrices. In addition, the baseline stated in the text is incorrect: with FoM = sqrt(det F) in two parameters, combining two independent probes with equal Fisher matrices doubles the FoM, not multiplies it by sqrt(2). The quantitative independence claim therefore needs to be rederived or softened.","section":"§7, Table 1"},{"comment":"The treatment of kNN tails is inconsistent. Section 5.3 defines the kNN CDFs and specifies the k values but no scale or tail cut, while Section 7 states that 'we have explicitly removed the tails of the kNN distributions in order to ensure that our summary statistics are Gaussian.' The threshold and the number of retained bins must be specified, and the impact of this cut on the Fisher results should be quantified; as written, the comparison protocol between kNNs and Betti curves is ambiguous.","section":"§5.3 vs §7"},{"comment":"The entire comparison rests on the assumption that each summary statistic is multivariate Gaussian with parameter-independent covariance, but the evidence for this is only the qualitative chi-square test in Appendix A. Given that the Betti curves are truncated specifically to enforce Gaussianity, and that the tails that are removed may contain non-Gaussian information, the paper should either provide a more quantitative Gaussianity assessment (for example, goodness-of-fit statistics or a comparison of Fisher information with and without tail bins) or explicitly frame the results as conditional on this assumption. The ranking of probes in Table 1 could change if the excluded tails carry information.","section":"§6, Eq. (3)"}],"minor_comments":[{"comment":"The abstract's statement that 'the kNNs provide very competitive constraints' overstates the results: in Table 1, DR-kNN has FoM 0.42 versus 0.52 for the 2-point function in the Omega_m-sigma8 space, so only DD-kNN is clearly competitive.","section":"Abstract"},{"comment":"The finding that beta2 does constrain n_s in the three-parameter space should be echoed in the abstract or conclusions to avoid overgeneralizing the 'beta2 has almost no constraining power' statement, which is specific to the reduced two-parameter space.","section":"§7, Fig. 9"},{"comment":"Please clarify whether the two-parameter Fisher matrices are recomputed from derivatives and covariances restricted to {Omega_m, sigma8} or derived by projecting the five-dimensional matrices; the eigenvectors in Table 2 are mixtures involving h and n_s, so the relationship between the 5D and 2D convergence statements is not immediate.","section":"§6.2, Table 2"},{"comment":"Please state the number of bins retained for each Betti curve after the beta_d > 5e-4 cut; Figure 3 suggests roughly 60 total bins, but the exact numbers matter for reproducibility and for interpreting the Appendix A chi-square degrees of freedom.","section":"§5.2"},{"comment":"Please label the number of degrees of freedom and the number of simulations used for each chi-square distribution in Figure 10, since the power of the Gaussianity test depends on both.","section":"Fig. 10"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a methods/comparison journal and contains a careful Fisher-convergence analysis using public simulations. The main concerns are that the headline claims (beta2 uninformative, kNNs and Betti curves not fully independent) rest on choices that are not fully tested or described: the Betti tail cut, the kNN tail treatment, and the construction of combined Fisher matrices. These issues are fixable within the manuscript's scope, so I would not reject, but they need to be addressed before the quantitative conclusions can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi, quick take on 2502.09709. The paper compares Betti curves and kNNs as cosmological summary statistics using simulation-based Fisher matrices from Quijote. The genuinely new piece is the convergence diagnostic: they diagonalize the Fisher matrix and fit each eigenvalue to F_inf + F_noise/N_deriv, which cleanly identifies which parameter directions are noise-dominated. That is a practical tool, and the care with Hartlap correction, 5000 covariance simulations, and restricted parameter space is evident. The comparison itself is also useful: DD-kNNs and Betti curves come out similar and better than 2-pt and DR-kNNs in {Omega_m, sigma_8}, and the claimed partial redundancy is a plausible, interesting result.\n\nThe main soft spot is the beta2 claim. The abstract says \"almost no constraining power comes from beta2,\" but this is parameter-dependent: in the 3-parameter space beta2 does constrain n_s better than the other Betti curves (Fig. 9). And the tail cut beta_d > 5e-4 (Sec. 5.2) removes exactly the large-r tail where beta2's signal may live, since beta2 decays to zero. The chi-square test in Appendix A is qualitative and has limited power for low-count bins, so excluding non-Gaussian tails may bias the comparison against beta2. The blanket statement should be softened or shown to be robust to the threshold.\n\nSecond: the FoM values in Table 1 have no uncertainty estimates. The convergence analysis addresses eigenvalue noise, but not the uncertainty on the FoM ranking itself. Differences of 10-20% could be within noise. The authors should at least discuss this, or ideally bootstrap.\n\nThird: Section 7 says kNN tails were explicitly removed, but Section 5.3 never specifies the cut. That is a real protocol ambiguity, and it matters for fair comparison. Also, no analysis code is released.\n\nNone of these sink the paper. The central comparison is careful and the convergence tool is a real contribution. But the beta2 result, as stated, is not robust to the tail treatment, and the FoM ranking is a point estimate. I would send this to peer review with a request for revisions: qualify the beta2 claim, add robustness tests around the Betti tail threshold, document the kNN cuts, and either release code or give uncertainty estimates on the FoM.","headline":"Careful Fisher comparison of Betti curves and kNNs with a useful convergence diagnostic; the beta2 claim is overstated and needs qualification, but the paper is worth refereeing.","tokens_in":16051,"tokens_out":3111,"would_cite":true,"duration_ms":31883,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["85A40","55N31"],"pacs":[],"model":"deepseek-v4-flash","headline":"For halo clustering on nonlinear scales, nearly all Betti-curve information sits in beta0 and beta1, while kNNs match it.","keywords":["Betti curves","k-nearest neighbor distributions","persistent homology","cosmological information","Fisher matrix convergence","Quijote simulations","halo clustering","non-Gaussian information"],"falsifier":"Recompute the Fisher comparison without the $\\beta_d > 5\\times10^{-4}$ truncation, or with a likelihood that allows non-Gaussian tails (for instance, simulation-based inference on the full data vectors), and see whether $\\beta_2$ or the $k$NN tails add constraints that change the probe ranking. Alternatively, increase $N_{\\mathrm{deriv}}$ well beyond 500 and check whether the third Fisher eigenvalue, aligned with $n_s$, converges; if it does, the reduced $\\{\\Omega_m, \\sigma_8\\}$ comparison may miss information that distinguishes the probes.","tokens_in":15035,"feed_emoji":"🔭","tokens_out":11561,"duration_ms":99038,"temperature":0.7,"pith_summary":"Using Fisher matrices built from 5,000 Quijote simulations, this paper sets out to compare how much cosmological information two modern summary statistics extract from halo clustering on scales $2\\,h^{-1}\\mathrm{Mpc} < r < 50\\,h^{-1}\\mathrm{Mpc}$: Betti curves, a persistent-homology summary of topological features, and $k$-nearest-neighbor ($k$NN) distance distributions. The authors find that the first two Betti curves, $\\beta_0$ and $\\beta_1$, which track connected components and loops, carry essentially all of the constraining power in the well-converged parameter space; the void-counting curve $\\beta_2$ contributes almost nothing to constraints on $\\Omega_m$ and $\\sigma_8$. They also find that $k$NNs, especially the data-data variant, give constraints competitive with the full Betti set, and that combining the two probes yields less information than independent statistics would. A secondary, methodological result is that only two parameter directions ($\\Omega_m$ and $\\sigma_8$) are reliably converged with the available simulations, which limits the dimensionality of the comparison.","feed_headline":"Betti curves and kNNs extract overlapping cosmology from halos","feed_subtitle":"Most topological information lives in beta0 and beta1; k-nearest neighbors match it.","key_machinery":"Betti curves are computed from an $\\alpha$-complex filtration of the halo point cloud: at each filtration scale $r$, $\\beta_d(r)$ counts the number of $d$-dimensional homology classes per halo (components for $d=0$, loops for $d=1$, voids for $d=2$). $k$NN CDFs record the fraction of query points whose distance to their $k$-th nearest halo is below $r$, with two variants: DR-$k$NNs (random volume-filling points to data) and DD-$k$NNs (data point to data point). The comparison instrument is the Fisher matrix from a Gaussian likelihood with constant covariance, with derivatives estimated by finite differences from 500 Quijote simulations; the paper's key methodological addition is an eigendecomposition-based convergence test that fits $F(N_{\\mathrm{deriv}}) = F_{\\infty} + F_{\\mathrm{noise}}/N_{\\mathrm{deriv}}$ for each eigenvalue and identifies which parameter directions are noise-dominated. This convergence test is what justifies restricting the headline comparison to the two well-constrained directions $\\Omega_m$ and $\\sigma_8$.","core_discovery":"The paper's central claim is that, for this halo sample, Betti curves and $k$NN distributions are measuring largely the same non-Gaussian clustering information, and that within the Betti curves the information is concentrated in the first two homology dimensions. Quantitatively, combining $\\beta_0$ and $\\beta_1$ gives 98% of the figure of merit of all three Betti curves in the $\\{\\Omega_m, \\sigma_8\\}$ plane; $\\beta_2$, the void-counting curve, gives weak constraints on $\\Omega_m$ and almost none on $\\sigma_8$. The DD-$k$NNs (distances between halo pairs) achieve a figure of merit nearly equal to the Betti curves, while the DR-$k$NNs (distances from random volume-filling points to halos) perform worse in this setup. Combining Betti curves with $k$NNs improves constraints, but not by the factor $\\sqrt{2}$ expected for independent probes, which the authors interpret as evidence that the two statistics are connected, possibly through their shared sensitivity to density regions.","pith_inferences":["The overlap between Betti curves and $k$NNs hints at a mathematical correspondence: the $\\alpha$-complex filtration threshold is itself a distance criterion, so the homology classes counted by $\\beta_d$ are functions of the same pairwise-distance information that $k$NN CDFs compress; one could test this by asking whether the Euler characteristic ($\\beta_0 - \\beta_1 + \\beta_2$) fully predicts the $","$\\beta_2$'s near-zero constraining power in the two-parameter plane may depend on the number-density cut at 150,000 halos; on sparser or denser samples, or with a void-oriented filtration, void counting could become informative, so the 'drop $\\beta_2$' advice should be re-checked per survey selection.","The DR-$k$NN versus DD-$k$NN split maps onto low-density versus high-density environments, which suggests the Betti curves' sensitivity could be decomposed the same way; a joint analysis in redshift space, where $k$NNs have a natural decomposition, could clarify which physical regions carry the shared signal.","A direct observational test would be to apply the same Fisher comparison to a real galaxy catalog in redshift space, where $k$NNs have known advantages; if the overlap persists, pipeline designers could choose $k$NNs on cost grounds alone."],"forward_implications":["If the result holds, future Betti-curve analyses on similar halo samples can drop $\\beta_2$ in the $\\Omega_m$-$\\sigma_8$ plane and lose little: $\\beta_0$ plus $\\beta_1$ recovers 98% of the combined figure of merit.","$k$-nearest-neighbor distributions, especially DD-$k$NNs, can serve as a computationally cheaper stand-in for Betti curves at these scales, with comparable constraining power and no need for $\\alpha$-complex triangulation.","Combining Betti curves with $k$NNs will not double the information; the realized gain is closer to the gain from adding $k$NNs to the two-point function, consistent with a shared non-Gaussian signal.","The 2-point correlation function adds almost no new information once Betti curves and both $k$NN variants are included, so the topological and neighbor statistics are absorbing the small-scale non-Gaussian information that the power spectrum misses.","Forecasts in higher-dimensional parameter spaces from Quijote-style derivative sets should first run the eigenvalue convergence test; the paper finds only two directions are reliably converged, with the $h$ and $M_\\nu$ directions fully noise-dominated and the $n_s$ direction still converging."],"supporting_citations":[{"why":"Supplies the Quijote simulation suite, halo catalogs, fiducial cosmology, and finite-difference derivative sets used for every Fisher matrix.","marker":"Villaescusa-Navarro et al. (2020)"},{"why":"Establishes the Betti-curve measurement pipeline, including the $\\alpha$-complex filtration and normalization, that this paper applies to halo catalogs.","marker":"Ouellette et al. (2023)"},{"why":"Introduces $k$NN CDFs as a cosmological summary statistic and proves their sensitivity to the full $n$-point hierarchy, the theoretical basis for the comparison.","marker":"Banerjee & Abel (2021a)"},{"why":"Defines and validates the DR- and DD-$k$NN variants used here and their different density-region sensitivities.","marker":"Yuan et al. (2023)"},{"why":"Supplies the correction factor applied to the inverse covariance matrix to unbias Fisher estimates from simulated covariances.","marker":"Hartlap et al. (2007)"},{"why":"Documents the noise bias in numerical-derivative Fisher matrices, motivating the eigendecomposition convergence check.","marker":"Coulton & Wandelt (2023)"},{"why":"Studies derivative noise in Quijote-based Fisher forecasts, providing the context for the paper's convergence analysis.","marker":"Wilson & Bean (2024)"},{"why":"Shows a connection between $k$NNs and Minkowski functionals, cited as support for the speculative link between $k$NNs and topological summaries.","marker":"Gangopadhyay et al. (2025)"}],"fun_headline_variants":["Betti curves and kNNs probe same halo clustering","Most cosmic info in beta0 and beta1 of Betti curves","kNNs match Betti curves for halo clustering constraints","Betti and kNN share overlapping cosmic information","Halo topology: beta0 and beta1 carry the signal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline comparison assumes each summary statistic follows a multivariate Gaussian distribution with a constant, parameter-independent covariance; the paper enforces this by truncating Betti curves at $\\beta_d > 5\\times10^{-4}$ and checks it qualitatively, but if the excluded tails carry non-Gaussian information, the ranking of probes could change.","fun_headline_variants_meta":{"raw":{"variants":["Betti curves and kNNs probe same halo clustering","Most cosmic info in beta0 and beta1 of Betti curves","kNNs match Betti curves for halo clustering constraints","Betti and kNN share overlapping cosmic information","Halo topology: beta0 and beta1 carry the signal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000554,"raw_usage":{"total_tokens":2687,"prompt_tokens":1037,"completion_tokens":1650,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":1582}},"tokens_in":653,"tokens_out":1650,"duration_ms":16140,"temperature":1.0,"reasoning_tokens":1582,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T20:45:03.840170+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Fisher comparison without the $\\beta_d > 5\\times10^{-4}$ truncation, or with a likelihood that allows non-Gaussian tails (for instance, simulation-based inference on the full data vectors), and see whether $\\beta_2$ or the $k$NN tails add constraints that change the probe ranking. Alternatively, increase $N_{\\mathrm{deriv}}$ well beyond 500 and check whether the third Fisher eigenvalue, aligned with $n_s$, converges; if it does, the reduced $\\{\\Omega_m, \\sigma_8\\}$ comparison may miss information that distinguishes the probes.","supporting_citations":[{"cited_title":"2023, MNRAS, 523, 5738, doi: 10.1093/mnras/stad1765","cited_arxiv_id":null,"evidence_quote":"Establishes the Betti-curve measurement pipeline, including the $\\alpha$-complex filtration and normalization, that this paper applies to halo catalogs."},{"cited_title":"Fisher's Mirage: Noise Tightening of Cosmological Constraints in Simulation-Based Inference","cited_arxiv_id":"2406.06067","evidence_quote":"Studies derivative noise in Quijote-based Fisher forecasts, providing the context for the paper's convergence analysis."}],"review_version":1}