{"id":"7b825e78-c868-4316-bac8-eac2550716e7","arxiv_id":"2507.06964","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new Julia package implements Ringbauer, Coop, and Barton (2017) IBD-based dispersal inference, adds numerical integration for arbitrary density models and automatic differentiation, and includes simulation and analysis pipelines.","lead":"This paper presents IdentityByDescentDispersal.jl, a Julia package that estimates effective population density and per-generation dispersal from shared IBD DNA blocks. It makes a 2017 statistical method easier to use, supports flexible demographic models through numerical integration, and ships with a simulation template and a Snakemake pipeline.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The only validation example's MLE for D is 282 with 95% CI 260–303, excluding Dtrue≈250, and Fig. 1's caption claims agreement. Replicate coverage is needed to tell whether this is bias, effective-density mismatch, or an unlucky draw before the package's core claim is trusted.","rationale":"The reader's weakest_assumption is about the external validity of the RCB17 model for real IBD data. I agree that is a limitation, but the more immediately load-bearing concern is internal: the paper's only validation example does not recover the simulated parameters within its reported uncertainty, and the Figure 1 caption overstates the agreement. This matters even if RCB17 is perfectly specified, because it casts doubt on the implementation itself or on the calibration of the reported intervals. The concrete replicate-coverage test I propose would settle whether the published numbers are an unlucky draw, an effective-density definitional issue, or an actual bug. I credit the package's reproducible assets—registered Julia package, SLiM template, and Snakemake pipeline—as real evidence that the software exists and runs; this is why I do not recommend rejection, only that the validation inconsistency be resolved before the correctness claim is accepted at face value. My recommendation therefore leaves the reader's CONDITIONAL verdict unchanged.","tokens_in":4413,"tokens_out":6274,"duration_ms":76419,"concrete_test":"Re-run the Section 4 example on at least 10 independent SLiM simulations with the same nominal constant-density parameters (D=250, σ=0.071) and error-free IBD detection, and report coverage of the 95% Fisher-information CIs and the Turing posterior credible intervals, along with the realized mean of each estimator. If coverage is approximately 95% and the estimator means are centered on the nominal parameters, the published example was an unlucky draw and the implementation is likely sound. If the CIs systematically exclude the nominal values, determine whether the bias reflects finite-sample effective density (by comparing against the realized coalescence density in the simulations) or a code-level error in Eq. (1) or the composite likelihood.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the package correctly implements Eq. (1) and can be used to estimate effective density D and dispersal σ. The only quantitative validation of this claim is Section 4: one constant-density, error-free simulation with Dtrue≈250 diploids/km² and σtrue≈0.071 km/generation. The reported MLE is D≈282 (95% CI 260–303) and σ≈0.068 (95% CI 0.065–0.071); the pseudo-posterior means are E[D|data]≈281 and E[σ|data]≈0.068. Thus the 95% CI for D excludes the simulated value, and the σ CI only touches its true value. This directly contradicts the Figure 1 caption, which says the pseudo-posterior 'is forced to concentrate near the true values.' The paper itself warns that IBD detection errors are common and that composite-likelihood inference is overconfident, yet the example uses error-free blocks and still fails to include the truth in the reported interval. The mismatch could mean (i) an implementation error in the expected-block density or likelihood, (ii) a mismatch between the nominal simulation parameters and the effective parameters estimated by the RCB17 model, or (iii) an unlucky finite-sample draw amplified by overconfident composite-likelihood intervals. The paper does not distinguish these possibilities, so the central 'it works' claim is not yet established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces IdentityByDescentDispersal.jl, a registered Julia package that implements the identity-by-descent (IBD) block inference scheme of Ringbauer, Coop, and Barton (2017). The package computes the expected density of IBD blocks per pair of individuals via Eq. (1), offers analytic solutions for constant and power-law effective density functions, supports arbitrary user-defined density functions through numerical quadrature, and provides a composite-likelihood objective that is compatible with automatic differentiation, enabling gradient-based maximum-likelihood and Bayesian inference with Turing.jl. The manuscript also describes a SLiM simulation template and a Snakemake pipeline for end-to-end analysis from phased VCF files to parameter estimates. A single validation example on simulated error-free IBD data reports MLEs of D≈282 diploids/km² and σ≈0.068 km/generation, which the authors interpret as success.","tokens_in":4617,"tokens_out":2714,"duration_ms":32441,"significance":"If the package works as claimed, it fills a clear gap: there is no general-purpose, maintained software implementation of the RCB17 method despite its demonstrated usefulness for estimating dispersal from spatial IBD data. The contributions are concrete and partly verifiable: the package is registered, the code is public, the AD-compatible composite likelihood is a genuinely useful engineering advance over the original gradient-free approach, and the extension to arbitrary density functions broadens the method's applicability. The accompanying simulation and bioinformatics pipelines are valuable for adoption. However, the manuscript's only quantitative validation is thin and internally inconsistent, so the central 'it recovers the true parameters' claim is not yet established. The strengths are real, but the evidence for the headline capability needs to be made rigorous before the package can be recommended for routine use.","major_comments":[{"comment":"The paper itself warns in Section 3 that composite-likelihood inference 'tend[s] to be overconfident, underestimating posterior uncertainty and yielding too narrow confidence intervals.' Yet the example uses Fisher-information-based 95% intervals and presents them as supporting the method. The reported interval for D excludes the simulated value, which is exactly the overconfidence failure mode the authors warn about. The manuscript does not address whether the discrepancy is due to the known overconfidence, to a bias in the expected-block density implementation, or to the effective-density interpretation of the SLiM parameters. This needs a concrete treatment: for example, report profile-likelihood intervals, run a small replicate study to show empirical coverage, or at least state explicitly that the intervals are known to be anti-conservative and should not be over-interpreted. As written, the validation does not support the uncertainty quantification that the example advertises.","section":"Section 3, last paragraph and Section 4"},{"comment":"A headline feature of the package is the ability to fit arbitrary user-defined density functions De(t) through numerical integration (Section 3, Table 1). However, no validation is shown that the numerical quadrature implementation produces results consistent with the analytic solutions for the constant-density and power-density cases. Since the custom-density path is central to the advertised novelty, the authors should add a convergence test or a comparison that demonstrates the numerical integral reproduces the analytic formulas to a stated tolerance over a range of parameter values. Without this, the claim that arbitrary demographic models are 'supported' is not quantitatively supported.","section":"Section 4, custom density feature"}],"minor_comments":[{"comment":"The paper title and running header use 'IdentityBydescentDispersal.jl' without a space between 'By' and 'Descent'; this should be corrected to 'IdentityByDescentDispersal.jl' for consistency with the package name and Julia conventions.","section":"Title and header"},{"comment":"Equation (1) is typed as 'G4t2 exp(−2Lt)· Φ(t|r,θ )dt', which lacks explicit multiplication signs and superscripts; it should read 'G · 4t² · exp(−2Lt) · Φ(t|r,θ) dt' (or use proper LaTeX). The current rendering is ambiguous and should be fixed for readability.","section":"Equation (1)"},{"comment":"The text says the package computes 'composite_loglikelihood' functions; the naming is consistent with the code, but the prose would benefit from a brief explanation of why 'composite' likelihood is used rather than full likelihood, especially for readers not familiar with RCB17.","section":"Section 3, software description"},{"comment":"The example code snippet uses 'contig_lengths = [1.0]', but the meaning of this value (genome length in Morgans?) is not explained in the text; a one-sentence clarification would help users interpret the pipeline output.","section":"Section 4, example"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a software/application note, and its central claim is that the package implements a known method correctly and provides a usable interface. The implementation appears consistent with the cited equations, and the code is public and registered, which is commendable. The main weakness is the validation: a single simulated dataset whose reported confidence intervals exclude the true D, combined with an over-optimistic caption, is not sufficient to establish the package's reliability. This is fixable within the manuscript's scope by adding a simulation-based calibration study (replicates, coverage, bias) or a careful discussion of effective-density interpretation and the known anti-conservatism of composite-likelihood intervals. I do not see a fatal flaw in the implementation itself, but the current evidence would not justify acceptance without the additional validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know one thing up front: this is a genuinely useful software paper, but the validation section is weaker than the contribution. If you need to estimate dispersal from IBD blocks, this package saves you from implementing RCB17 from scratch, and the AD-compatible composite likelihood plus numerical quadrature for arbitrary De(t) are real extensions. The authors are scrupulous about crediting the theory to Ringbauer et al. 2017, and the package is registered, documented, and ships with a SLiM template and a Snakemake pipeline. That is solid engineering and I expect the package to get used.\n\nThe soft spot is exactly what the stress-test note flags. There is only one validation dataset: error-free blocks from a constant-density simulation with Dtrue≈250 and σtrue≈0.071. The reported MLE is D≈282 with a 95% CI of 260–303, which excludes the true value, and the CI for σ only just touches its true value. The Figure 1 caption says the pseudo-posterior “is forced to concentrate near the true values,” but the numbers in the text contradict that. The authors even warn that composite-likelihood inference is overconfident, yet they don't discuss why their own interval misses the truth. That mismatch could be an implementation bug, a census-versus-effective density issue, or an unlucky draw amplified by overconfident intervals. The paper doesn't distinguish these.\n\nI don't think this is fatal. The code is transparent, the equations match RCB17, and the extension to custom density models is credible. But the central claim “this implementation recovers D and σ under a known demographic model” is not yet established. A revision with replicate simulations, coverage checks, and some honest discussion of effective-density interpretation would settle it. This is exactly the kind of paper where a serious referee can ask for a small simulation study and the authors can deliver it quickly.\n\nVerdict: send it to peer review. It deserves referee time, not a desk reject, but the referee should push on the validation.\n\nNet: I'd cite this if I were doing IBD-based dispersal inference, and I'd bring it to a methods reading group to discuss the validation standard for software papers. Just don't take the Figure 1 caption at face value until they fix the numbers.","headline":"Useful Julia package implementing RCB17 with a real extension to arbitrary demography, but its single validation example overclaims: the reported 95% CI excludes the true density, so the package's core inference claim needs more evidence.","tokens_in":5226,"tokens_out":1614,"would_cite":true,"duration_ms":19715,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents IdentityByDescentDispersal.jl, a registered Julia package that implements the Ringbauer–Coop–Barton identity-by-descent inference scheme, letting researchers estimate effective population density and per-generation…","keywords":["identity-by-descent","dispersal rate","effective population density","composite likelihood","Julia","isolation by distance","population genetics"],"falsifier":"Run many simulated populations with known $D$ and $\\sigma$ through the package's own SLiM-to-inference pipeline and measure the empirical coverage of the reported 95% confidence intervals; if the intervals contain the true values less than 95% of the time even when the generative model matches the assumed Poisson/coalescent model, the central inference claim fails.","tokens_in":4135,"feed_emoji":"🧬","tokens_out":11589,"duration_ms":109349,"temperature":0.7,"pith_summary":"This paper presents IdentityByDescentDispersal.jl, a registered Julia package that turns patterns of shared identity-by-descent (IBD) DNA blocks into estimates of effective population density and per-generation dispersal rate. It fills the stated gap of a general-purpose software implementation of the inference scheme proposed by Ringbauer, Coop, and Barton (2017), which had previously been limited to analytically tractable power-law density histories and gradient-free optimization. The package solves the core integral for the expected density of IBD blocks, evaluates composite likelihoods under a Poisson model, and supports maximum-likelihood and Bayesian inference. It also accepts arbitrary user-defined density functions through numerical integration and provides a simulation template plus an end-to-end pipeline from phased VCF files to estimates. If the implementation is right, researchers studying fragmented or invasive populations can obtain dispersal estimates from spatial genetic data without writing the method from scratch.","feed_headline":"Shared DNA blocks now yield dispersal estimates in one package","feed_subtitle":"Population density and per-generation dispersal follow from phased genomes and pairwise distances.","key_machinery":"The load-bearing object is Eq. (1), the expected density of IBD blocks: $$\\mathbb{E}[N_L \\mid r,\\$\\theta$] = \\int_0^\\infty G\\,$4t^{2}$ $e^{{-2Lt}}$\\,\\Phi(t\\mid r,\\$\\theta$)\\,dt,$$ where $G$ is the genome length in Morgans, $L$ is block length, $r$ is the present-day distance between the two individuals, and $\\Phi(t\\mid r,\\theta)$ is the instantaneous coalescence rate at time $t$ under the demographic model with parameters $\\theta$. The package evaluates this integral either analytically for a constant density or a power-law density $D_e(t)=Dt^{-\\beta}$, or numerically through Gaussian quadrature for a user-defined $D_e(t)$. It then forms a composite log-likelihood by assuming the number of IBD blocks in a small length bin $[L,L+\\Delta L]$ is Poisson with mean $\\mathbb{E}[N_L\\mid r,\\theta]\\Delta L$. Automatic differentiation of this composite likelihood is what makes gradient-based optimization and Bayesian sampling practical.","core_discovery":"On the paper's own terms, the central claim is that the identity-by-descent inference scheme of Ringbauer, Coop, and Barton (2017) can be implemented as a general-purpose, registered Julia package and made practical for ordinary users. The package computes the expected density of IBD blocks per pair of individuals and per unit block length from Eq. (1), forms a composite log-likelihood under the Poisson assumption, and delivers MLEs or Bayesian pseudo-posteriors for the effective population density $D$ and dispersal rate $\\sigma$. It extends the original method by allowing arbitrary user-defined density histories $D_e(t)$ through numerical integration and by making the composite likelihood fully compatible with automatic differentiation, including differentiation with respect to the power-law exponent $\\beta$. For a simulated constant-density population with true values $D\\approx 250$ diploids/km$^2$ and $\\sigma\\approx 0.071$ km/generation, the worked example reports $D_{\\mathrm{MLE}}\\approx 282$ diploids/km$^2$ (95% CI 260\\textendash 303) and $\\sigma_{\\mathrm{MLE}}\\approx 0.068$ km/generation (95% CI 0.065\\textendash 0.071), alongside a Bayesian pseudo-posterior concentrated near the true values.","pith_inferences":["A natural next step would be to benchmark the package on real datasets with known demographic histories, such as island populations or invasion fronts, since the paper only demonstrates a simulated example.","Because the expected-block calculation is modular in $\\Phi(t\\mid r,\\theta)$, the same code could be adapted to heterogeneous or anisotropic landscapes without rederiving the integral, although the paper does not propose this.","The paper's own caution that composite-likelihood intervals are overconfident implies that conservation decisions should treat the reported confidence intervals as lower bounds on uncertainty until simulation-based calibration is performed."],"forward_implications":["Users can obtain composite-likelihood MLEs or Bayesian posteriors for $D$ and $\\sigma$ from IBD data using a registered package rather than reimplementing the underlying equations.","Arbitrary histories of the effective density $D_e(t)$ become estimable numerically, not just the analytically tractable power-law form $Dt^{-\\beta}$.","Gradient-based optimization and modern MCMC samplers work directly on the composite likelihood because it is compatible with automatic differentiation, including with respect to $\\beta$.","The supplied SLiM template and Snakemake workflow take users from phased VCF files to final estimates, making the method accessible in applied conservation and evolutionary studies."],"supporting_citations":[{"why":"Supplies the core inference scheme: Eq. (1) for the expected density of IBD blocks and the Poisson composite-likelihood framework the package implements.","marker":"Ringbauer, Coop, and Barton 2017"},{"why":"Provides the Gauss–Kronrod quadrature used to evaluate Eq. (1) for arbitrary user-defined density functions $D_e(t)$.","marker":"Johnson 2013"},{"why":"Underpins the automatic-differentiation-compatible formulation of the composite likelihood, enabling gradient-based optimization.","marker":"Geoga et al. 2022"},{"why":"Turing.jl is the probabilistic programming framework used in the example to run Bayesian inference on the composite likelihood.","marker":"Ge, Xu, and Ghahramani 2018"},{"why":"HapIBD detects IBD blocks from phased VCF data in the package's end-to-end pipeline.","marker":"Zhou, Browning, and Browning 2020"},{"why":"Refined IBD post-processes HapIBD output before the preprocessed data are passed to the likelihood functions.","marker":"B. L. Browning and Browning 2013"},{"why":"SLiM and tree-sequence recording power the simulation template used to generate synthetic datasets for calibration.","marker":"Haller and Messer 2023; Haller et al. 2019"},{"why":"Snakemake structures the bioinformatics pipeline from phased VCF files to final estimates.","marker":"Mölder et al. 2021"}],"fun_headline_variants":["IBD blocks now estimate density and dispersal","From IBD blocks to dispersal and density in Julia","New Julia tool infers density and dispersal from IBD","IBD blocks yield density and dispersal estimates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole method rests on the assumption that the mathematical model behind Eq. (1), together with the Poisson composite likelihood, correctly describes how IBD blocks appear in real genomes, so a wrong generative model would bias the estimates even if the code runs perfectly.","fun_headline_variants_meta":{"raw":{"variants":["IBD blocks now estimate density and dispersal","From IBD blocks to dispersal and density in Julia","New Julia tool infers density and dispersal from IBD","IBD blocks yield density and dispersal estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000915,"raw_usage":{"total_tokens":3935,"prompt_tokens":956,"completion_tokens":2979,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":2921}},"tokens_in":572,"tokens_out":2979,"duration_ms":21079,"temperature":1.0,"reasoning_tokens":2921,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:50:37.527689+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run many simulated populations with known $D$ and $\\sigma$ through the package's own SLiM-to-inference pipeline and measure the empirical coverage of the reported 95% confidence intervals; if the intervals contain the true values less than 95% of the time even when the generative model matches the assumed Poisson/coalescent model, the central inference claim fails.","supporting_citations":[],"review_version":1}