{"id":"a36c3a31-9bb9-4f19-8ec5-31994e420b99","arxiv_id":"1908.08608","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A pair-site composite likelihood model separately estimates fixed gene conversion tract lengths and initiation rates, yielding short exonic and long intronic tract estimates.","lead":"This paper introduces a composite likelihood method to estimate the rate and average length of fixed interlocus gene conversion events from DNA alignments of duplicated genes. It applies the method to yeast, primate protein-coding, and primate intron data, finding that gene conversion accounts for a substantial share of evolutionary change.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulation and bootstrap protocols fix τ and estimate only p, so the claimed separate estimation of tract length and initiation rate is never actually exercised.","rationale":"The reader identified the divergence-independence assumption as the weakest link. That is an acknowledged biological limitation and could bias tract-length estimates, but it is a modeling assumption external to the inference machinery. The more immediate, internal problem is that the numerical experiments never estimate the two quantities that the method claims to separate: every simulation and bootstrap analysis fixes τ and optimizes only p. This means the central methodological claim is supported by an indirect consistency check rather than by exercising the actual estimation procedure. The paper's own caveat in §5.3 about underestimated uncertainty reinforces the concern. Because the broad empirical conclusion that IGC is substantial likely survives this issue, the appropriate verdict remains conditional: the method is promising and the qualitative conclusion is plausible, but the quantitative separate estimates of tract length and initiation rate require validation with joint estimation before they are taken at face value. The reader's verdict of CONDITIONAL is therefore unchanged, though for a different reason than the one emphasized in the reader's weakest-assumption analysis.","tokens_in":15366,"tokens_out":5395,"duration_ms":62006,"concrete_test":"Re-run the simulation experiment of §5.2 with full joint estimation: for each simulated alignment, maximize the composite likelihood over p and η (or p and τ) simultaneously, starting from multiple initial values, and report the joint sampling distribution and profile likelihood for (η,p). Also construct a 95% percentile interval for 1/p from these joint estimates and compare coverage with the conditional interval of §5.3. If the joint estimates show a pronounced ridge or poor coverage, the separate-estimation claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central methodological claim is that the PS approach separately estimates the fixed IGC tract-length distribution (via 1/p) and the IGC initiation rate η. This claim is not tested by the paper's own validation. In the simulation study (§5.2), the composite likelihood is maximized only over p, with τ = η/p and all other parameters fixed at their simulation values; η is then recovered indirectly as τ/(1/p). In the parametric bootstrap (§5.3), each replicate is analyzed by estimating p while constraining all other free parameters, including τ, to the point estimates from the original data. Thus neither the simulations nor the bootstrap exercise the claimed two-parameter separation. The reported bootstrap intervals for 1/p are conditional on a fixed τ and therefore do not include uncertainty in τ, branch lengths, or rate parameters; the paper itself acknowledges that this 'will presumably cause the uncertainty in the tract length parameter p to be underestimated' (§5.3). Because τ and p are linked through τ = η/p, a ridge or weak identifiability in the (η,p) plane would not be detected by this protocol. The empirical short exonic tract-length estimates in Table 2 are consequently reported with uncertainty intervals that may be artificially narrow, and the central claim that the two quantities are separately estimable is supported only by an indirect consistency check, not by the full inference procedure.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a pair-site (PS) composite-likelihood approach to estimating the parameters of interlocus gene conversion (IGC) from alignments of duplicated genes. The model extends the authors' earlier independent-site (IS) model by jointly considering pairs of sites in two paralogs, allowing a geometric distribution for the length of fixed IGC tracts (parameter p) and an initiation rate η, with τ = η/p. The authors derive the resulting 256-state rate matrix for pairs of sites, apply the method to simulated data and to yeast paralogs, primate EDN/ECP genes, and primate introns, and report estimates of mean fixed tract length (1/p) and the proportion of substitutions attributable to IGC. They conclude that IGC accounts for a substantial fraction of evolutionary change in duplicated regions, while acknowledging that the tract-length estimates from protein-coding data are short and uncertain.","tokens_in":15617,"tokens_out":3315,"duration_ms":35062,"significance":"If the central claim that the PS approach can separately estimate IGC tract-length distribution and initiation rate is supported, the paper provides a useful methodological advance over the IS model and over the binary-state approach of Harpak et al. (2017), because it works with full nucleotide states and does not require the HMM assumption that each site experiences at most one IGC event. The rate-matrix construction is explicit and the composite-likelihood framework is clearly connected to McVean et al. (2002). The paper also includes a valuable external control, the paralog-swapping experiment, which supports the reality of the detected IGC signal. Software and data are made available. However, the validation as presented does not actually exercise the claimed two-parameter separation, and several empirical estimates rest on acknowledged but potentially strong assumptions.","major_comments":[{"comment":"The validation protocol never estimates η and p jointly. In the simulation study, the composite likelihood is optimized only over p with τ fixed at the value used to simulate; in the parametric bootstrap, only p is estimated with all other parameters constrained to point estimates, in particular τ. Therefore the simulations and bootstrap do not test the central claim that the PS approach can separately estimate the tract-length parameter p and the initiation rate η. Because τ = η/p, any ridge or weak identifiability in the (η,p) plane would remain undetected, and the bootstrap intervals for 1/p reported in Table 2 are conditional on an estimated τ. I ask the authors to run full joint inference in at least a subset of simulation scenarios, or to provide a profile composite-likelihood or Godambe-information analysis demonstrating that (η,p) are separately identifiable.","section":"§5.2 and §5.3"},{"comment":"The simulation summary discards estimates of 1/p that exceed the true value by a factor of 10 or more before computing sample means, while including those discarded values in the reported interquartile ranges. This asymmetric treatment can make the point estimates look better than the procedure actually is, especially at true lengths 200–500 where 3, 7, 6, and 7 estimates were discarded. Please report the full distribution of estimates, the proportion discarded under each condition, and whether the conclusions change when no values are discarded. If long-tailed estimates are inherent to the estimator, that property should be stated as a characteristic of the method rather than hidden by the mean calculation.","section":"§3.1 and Figure 1"},{"comment":"The intron analysis is reported with parameter values shared across the five groups of introns, which the authors explicitly equate to assuming that the duplications generating all five groups occurred at the same time. The text also states that separately analyzing the groups gave substantially different maximum composite-likelihood estimates from different starting values, indicating numerical instability. The intron estimate of 1/p ≈ 370 and its bootstrap interval therefore rest on an assumption known to be violated and on an optimization landscape that is not stable. This limitation should be stated in the results or abstract, or the sensitivity of the intron estimate to the shared-duplication-time assumption should be quantified.","section":"§5.6 and Table 2"},{"comment":"The model assumes that IGC mutations occur and fix at a rate independent of the sequence divergence between donor and recipient paralogs, and the authors cite Chen et al. (2007) as evidence against this assumption. If IGC rates decline as paralogs diverge, the model could misattribute divergence-dependent declines in gene conversion to short tract lengths or low initiation rates, directly affecting the central estimates of p and η. A concrete check would be to fit a variant of the PS model with a divergence-dependent IGC rate, or to compare estimates on branch subsets with different divergence levels. Without such a check, the empirical tract-length estimates remain vulnerable to this known biological violation.","section":"§4.1 and §2.1"}],"minor_comments":[{"comment":"In the parametric bootstrap description, the phrase 'constraining all other free parameters to their true values' is inaccurate: bootstrap replicates are simulated from the original point estimates, not from true values. Please say 'estimated values' or 'point estimates'.","section":"§5.3"},{"comment":"The caption refers to the 'maximum log-likelihood value' of the HKY+IS-IGC model, but the procedure maximizes a composite likelihood, not a true likelihood. Please use 'maximum composite log-likelihood' to avoid confusion.","section":"Table 1 caption"},{"comment":"The text says 'the actual YDR418W YER054C data has a gap' but the gene name is YEL054C everywhere else in the paper (e.g., Tables 1 and 2). Please correct this typo.","section":"§5.2"},{"comment":"The description of the intron filtering could be clearer about whether the 20 intron subsets are spread across 5 gene pairs with 4 introns each, or whether the distribution is uneven. A sentence specifying the number of introns per gene pair would help readers interpret the shared-parameter analysis.","section":"§5.6"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper extends Ji and Thorne's earlier independent-site IGC model to a pair-site composite likelihood that is meant to disentangle fixed IGC tract length (via 1/p) from the initiation rate (η). That is a real extension, and the model derivation is clear. They also add codon-position rate heterogeneity to the pairwise model, which is new relative to Harpak et al.'s binary treatment. Credit is due: the paralog-swapping control is a sensible check, the authors are explicit about assumptions they know are wrong, and the data and code are available.\n\nThe empirical bottom line – that IGC is responsible for a substantial fraction of change in duplicated regions – is probably right. The τ estimates are consistent across the IS and PS models, and the intron estimates of tract length are in a reasonable range. The short exonic tracts are the interesting result, but also the one I would not trust yet.\n\nThe soft spot is validation. In the simulations, the composite likelihood is maximized only over p, with τ fixed at the true value. The parametric bootstrap does the same, fixing τ at the point estimate. So the paper never actually demonstrates that the approach can estimate tract length and initiation rate separately. The reported interquartile ranges for 1/p are conditional on τ and are therefore too narrow, as the authors admit. Because τ = η/p, a weak identifiability ridge in (η, p) would not be detected by these exercises.\n\nThe other issue is the acknowledged assumption that IGC rate is independent of paralog divergence. This is false (Chen et al. 2007), and it could push the model to attribute divergence-dependent decline to short tracts or low initiation. The authors flag this, but it remains a real bias vector.\n\nOverall: a serious paper with an honest limitation section. The method is plausible but the central two-parameter separation is not yet supported by the evidence presented. The broad IGC result is secure enough to act on. I would send this to peer review with the request that simulations and bootstrap vary τ and p jointly, and that the divergence-dependence issue be addressed at least by sensitivity analysis.\n\nWorth a look if you work on multigene families or composite likelihood phylogenetics.","headline":"The paper's honest about its own limits, but the central two-parameter claim is never actually tested; the robust part is that IGC is common, not that tract lengths are short.","tokens_in":16105,"tokens_out":2281,"would_cite":false,"duration_ms":23320,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Estimator splits gene-conversion tract length from initiation rate","keywords":["interlocus gene conversion","composite likelihood","tract length","initiation rate","multigene family evolution","phylogenetic inference","gene duplication","nucleotide substitution model"],"falsifier":"Simulate alignments under a model where the IGC initiation rate declines with donor-recipient divergence while the true mean tract length is held constant, then fit the PS model; if the estimated mean tract length shrinks as divergence increases, the claimed separation of tract length from initiation rate is refuted.","tokens_in":15152,"feed_emoji":"🧬","tokens_out":5188,"duration_ms":48101,"temperature":0.7,"pith_summary":"Interlocus gene conversion (IGC) copies a stretch of DNA from one duplicated gene to its partner, erasing substitution evidence and making duplicated genes look more similar than their true history warrants. This paper asks whether the observable homogenization can be separated into two hidden quantities: how often fixed IGC events start and how long the fixed tracts are. It claims the answer is yes, via a pair-site (PS) composite likelihood that models two nucleotide sites in two paralogs jointly and attributes correlated change to a geometric tract-length distribution. Applied to yeast protein-coding paralogs, primate ribonuclease genes, and primate introns, the procedure estimates that IGC accounts for a substantial share of evolutionary change in duplicated regions. The paper also finds that estimated fixed tracts in protein-coding exons are unexpectedly short and uncertain, whereas intron estimates fall near the range seen in studies of IGC mutations.","feed_headline":"Estimator splits gene-conversion tract length from initiation rate","feed_subtitle":"Pairwise likelihood on duplicated genes recovers mean tract length and initiation rate, showing IGC drives a large share of their evolution.","key_machinery":"The pair-site (PS) composite likelihood: for two paralogs, the state of two sites in each paralog is a four-nucleotide joint state, giving a 256-state rate matrix whose off-diagonal entries add point-mutation rates and IGC rates that depend on the separation $n$ between sites. The IGC contribution decomposes as $\\eta/p$ for events covering a site, $\\eta/p [1-(1-p)^n]$ for events covering one of two sites but not the other, and $(\\eta/p)(1-p)^n$ for events covering both; the geometric parameter $p$ thus controls how quickly IGC-induced correlation decays with sequence distance, which is what lets the procedure estimate tract length separately from initiation rate. This is a pairwise composite likelihood in the style of recombination-rate estimation, with uncertainty assessed by parametric bootstrap.","core_discovery":"The central discovery is a statistical separation that previous phylogenetic IGC models could not achieve: the product of initiation rate and mean tract length ($\\tau = \\eta/p$) is replaced by two identifiable parameters, the fixed-tract initiation rate $\\eta$ and the geometric tract-length parameter $p$, with mean fixed tract length $1/p$. The separation works by extending the independent-site (IS) model to pairs of sites: for two paralogs and two aligned positions separated by $n$ sites, IGC events that cover both positions occur at rate $(\\eta/p)(1-p)^n$, and events covering only one position at rate $(\\eta/p)[1-(1-p)^n]$, so the distance-dependence of the correlation carries the tract-length signal. The model is fitted by maximum composite likelihood over all pairs of alignment columns, using Felsenstein's pruning algorithm for each pairwise marginal likelihood. Analyses of simulated data show average estimates of mean tract length close to truth under short-tract conditions, with growing variability as true tracts lengthen. On real data, the procedure attributes roughly 9 to 43 percent of nucleotide change to IGC across datasets and yields estimated mean fixed tract lengths of about 1 to 100 nucleotides for exons and about 370 nucleotides for introns.","pith_inferences":["If the rate of IGC truly declines as paralogs diverge, current PS estimates of short exonic tracts may be inflated artifacts; a natural extension is a model where $\\eta$ depends on divergence, which would let the procedure distinguish 'conversion stopped early' from 'conversion happened but was overwritten'.","The same pair-site machinery could be applied to more than two paralogs by summing over pairs, although the paper leaves this as future work.","A direct test of the exon-intron tract-length gap: analyze multi-exon protein-coding genes with long introns, where a single IGC event spanning an intron would create a distinctive correlation pattern across exon boundaries.","Because the composite likelihood ignores correlation beyond pairs, the variance of tract-length estimates is likely understated; block-bootstrap or full-likelihood checks on short alignments would quantify this."],"forward_implications":["If the PS separation is correct, estimates of $\\tau$ alone are insufficient: evolutionary analyses of duplicated genes should report or model tract length and initiation rate separately.","IGC should be included in studies of multigene family evolution, because ignoring it misattributes a substantial fraction of change to point substitution.","Protein-coding exons may have fixed IGC tracts far shorter than IGC mutation tracts, implying that selection or subsequent homologous recombination trims converted segments before fixation.","Primate intronic IGC tracts are long, on the order of hundreds of nucleotides, in line with mutation-level studies.","The near-equality of $\\tau$ estimates between IS and PS models indicates that the simpler IS model captures the overall IGC rate, while tract length information requires the pair-site extension."],"supporting_citations":[{"why":"Supplies the independent-site model and the yeast paralog datasets that the pair-site extension builds on.","marker":"Ji et al. (2016)"},{"why":"Provides the pairwise composite-likelihood template for estimating rates from pairs of sites.","marker":"McVean et al. (2002)"},{"why":"Gives the pruning algorithm used to compute each pairwise marginal likelihood.","marker":"Felsenstein (1981)"},{"why":"Independently developed a similar pair-site IGC procedure and supplied the primate intron alignments reanalyzed here.","marker":"Harpak et al. (2017)"},{"why":"Supplies the EDN/ECP sequences and tree used for the primate protein-coding analysis.","marker":"Zhang et al. (1998)"},{"why":"Cited as evidence that IGC becomes less likely as paralogs diverge, the assumption the model deliberately violates.","marker":"Chen et al. (2007)"},{"why":"Justifies the asymptotic unbiasedness of composite likelihood estimates and the parametric-bootstrap uncertainty approach.","marker":"Varin et al. (2011)"},{"why":"Shows IGC can oppose neofunctionalization, framing why the estimated IGC abundance matters.","marker":"Teshima and Innan (2008)"}],"fun_headline_variants":["Pairwise likelihood separates IGC tract length from initiation rate","Composite likelihood disentangles gene conversion length and rate","New model splits gene conversion tract length from initiation rate","IGC tract length and initiation rate now separately estimable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model assumes that gene conversion events occur and become fixed at a rate that does not depend on how different the two paralogs already are; if conversion actually slows as paralogs diverge, the procedure can mistake that slowdown for short tracts or low initiation rates.","fun_headline_variants_meta":{"raw":{"variants":["Pairwise likelihood separates IGC tract length from initiation rate","Composite likelihood disentangles gene conversion length and rate","New model splits gene conversion tract length from initiation rate","IGC tract length and initiation rate now separately estimable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1598,"prompt_tokens":990,"completion_tokens":608,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":544}},"tokens_in":606,"tokens_out":608,"duration_ms":6495,"temperature":1.0,"reasoning_tokens":544,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:34:09.944204+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate alignments under a model where the IGC initiation rate declines with donor-recipient divergence while the true mean tract length is held constant, then fit the PS model; if the estimated mean tract length shrinks as divergence increases, the claimed separation of tract length from initiation rate is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the pairwise composite-likelihood template for estimating rates from pairs of sites."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the pruning algorithm used to compute each pairwise marginal likelihood."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Independently developed a similar pair-site IGC procedure and supplied the primate intron alignments reanalyzed here."},{"cited_title":"EDN/ECP” row represents the primate EDN/ECP data set. The “Introns","cited_arxiv_id":null,"evidence_quote":"Supplies the EDN/ECP sequences and tree used for the primate protein-coding analysis."},{"cited_title":"N., Chuzhanova, N., F´ erec, C., and Patrinos, G","cited_arxiv_id":null,"evidence_quote":"Cited as evidence that IGC becomes less likely as paralogs diverge, the assumption the model deliberately violates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies the asymptotic unbiasedness of composite likelihood estimates and the parametric-bootstrap uncertainty approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows IGC can oppose neofunctionalization, framing why the estimated IGC abundance matters."}],"review_version":1}