{"id":"712a88cc-9314-489b-88ea-dd3037e2afba","arxiv_id":"2508.11461","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"An importance-sampling estimator approximates marginal sequence likelihoods for context-dependent evolution models with matching finite-sample error bounds and complexity that scales with the number of observed mutations rather than sequence length.","lead":"This paper proposes an importance-sampling algorithm to estimate the likelihood of molecular sequence evolution when mutation rates at each DNA site depend on neighboring sites, a setting where exact computation is usually intractable. If the claimed error and complexity bounds hold, it would make likelihood-based analyses of context-dependent evolution practical for long sequences.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central complexity-in-r claim is conditioned on an undefined 'practical regime'; the supplied text provides no proposal or variance proof to check.","rationale":"The reader's verdict was UNVERDICTED because the full text is corrupted, and I likewise cannot verify the proof or proposal. The reader's weakest assumption points to the undefined 'practical regime' and proposal/variance control; that is the same substantive concern I identify. I do not see a concrete technical flaw beyond the inability to inspect the argument, so my concern does not move the verdict; it reinforces the existing 'unverified' status. If a readable version confirms the regime and the variance bound, the central claim may be sound, but until then the complexity-in-r claim remains unsupported.","tokens_in":22885,"tokens_out":3919,"duration_ms":48349,"concrete_test":"Obtain a readable full text from arXiv. Locate the theorem defining the 'practical regime' and stating the finite-sample bounds. Check the proposal distribution, and compute the second moment of the importance weights for a simple context-dependent model (e.g., CpG→TpG hypermutability) at a realistic r/n. If the variance is infinite or the regime excludes realistic r/n, the matching-order complexity claim is vacuous or false.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim—finite-sample matching upper/lower error bounds with complexity growing in r rather than exponentially in n—depends on two unstated conditions: (i) a precise definition of the 'practical regimes of r/n' within which the bound holds, and (ii) a specification of the importance proposal and a proof that its variance is controlled in those regimes. The full text supplied is unreadable mojibake, with an unrelated arXiv identifier (2508.11468v2 [cs.SE]) embedded, so neither the theorem statement nor the proposal construction is inspectable. This is load-bearing because importance sampling is only useful when the proposal overlaps the target conditional distribution over mutational histories; a poor proposal yields an unbiased but potentially infinite-variance estimator, so the claimed O(r) complexity would not follow. The 'practical regime' qualifier could be vacuous: if the constant/prefactor depends on n outside a narrow band of r/n, the headline claim would not cover typical phylogenetics data (e.g., n ~ 10^4, r ~ 10^2).","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a randomized importance-sampling estimator for the marginal likelihood of two aligned molecular sequences under context-dependent (site-dependent) substitution models. The abstract claims finite-sample upper and lower bounds on the approximation error that match in order, and that, for 'practical regimes of r/n' (with r observed mutations and sequence length n), the sampler's complexity grows in r rather than exponentially in n. It further claims problem-specific complexity bounds for a named dependent-site model from the phylogenetics literature. The abstract is internally coherent and appropriately hedged, and no circular reasoning or fitting-to-data is apparent. However, the supplied full text is undecodable mojibake: every section, equation, and proof is unreadable, and an unrelated arXiv identifier (2508.11468v2 [cs.SE]) appears embedded in the text. Consequently, none of the paper's central claims can be independently checked from the submitted material.","tokens_in":22996,"tokens_out":4199,"duration_ms":53663,"significance":"If the claimed results hold, they would be a substantive contribution to computational phylogenetics. Context-dependent models have state spaces exponential in context length, and the marginal likelihood under such models is a recognized computational bottleneck. Matching-order finite-sample error bounds are stronger than a mere asymptotic consistency statement, and a complexity-in-r result would be practically valuable in the sparse-mutation regime typical of many comparative genomics problems. The abstract's contribution is clearly an estimator plus error bounds, not a parameter fit, so there is no obvious circularity. That said, the manuscript as supplied contains no readable derivation, no proposal construction, no theorem statement, and no variance argument; the strengths are claims about a result, not a verifiable result.","major_comments":[{"comment":"The body text supplied for review is unreadable mojibake. No definition, theorem, algorithm, or proof can be reconstructed. The embedded string 'arXiv:2508.11468v2 [cs.SE]' is an unrelated identifier and further indicates that this is not the paper's actual body. This is load-bearing: the central claim is a mathematical guarantee with matching bounds, and without an inspectable proof there is no basis for a soundness assessment.","section":"Full Text (throughout)"},{"comment":"The phrase 'for practical regimes of r/n' is never defined, and the full text does not clarify it. The headline complexity-in-r claim is conditional on this regime. The authors must state precisely what the regime is and how all hidden constants depend on n, r, and model parameters. If the regime excludes realistic data—for example, n ~ 10^4 with r ~ 10^2—the main claim would be vacuous for the applications the paper cites.","section":"Abstract"},{"comment":"The importance proposal is not described anywhere in the supplied text. Classical importance sampling is unbiased but can have infinite variance if the proposal and the target conditional distribution over mutational histories have poor overlap. The claimed finite-sample error bounds must rest on an explicit control of the proposal's variance. The paper needs to specify the proposal and prove the second-moment or concentration condition that the bounds use.","section":"Abstract (proposal construction omitted)"}],"minor_comments":[{"comment":"The notation r/n is used without first stating that r is the number of observed mutations and n is the sequence length; this should be stated explicitly in the abstract.","section":"Abstract"},{"comment":"The corrupted text and the unrelated arXiv identifier suggest a PDF generation or encoding failure. The authors should ensure that the submitted manuscript is a clean, readable PDF before resubmission.","section":"Full Text"}],"recommendation":"major_revision","confidential_remarks":"The paper's topic is relevant for stat.CO and computational phylogenetics. The abstract suggests a potentially interesting theoretical result, but the provided full text is not reviewable. I would ask the editor to obtain a clean version before any substantive review; if the mojibake is actually the arXiv submission, the manuscript should be returned for a corrected file. I am not recommending rejection because a readable resubmission could in principle address all concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a plausible theory paper that, if correct, gives a genuinely useful importance-sampling estimator for marginal likelihoods in context-dependent sequence evolution, with matching order finite-sample upper and lower bounds and complexity growing in r rather than exponentially in n. That's a meaningful claim for phylogenetics, where these likelihoods are a real bottleneck. The abstract is well-written and appropriately hedged with 'for practical regimes of r/n,' and the plan to specialize the bounds to a specific dependent-site model is a good move.\n\nBut I cannot vouch for any of it. The full text supplied to me is unreadable mojibake, with an unrelated arXiv identifier (2508.11468v2, cs.SE) embedded in the body. So no theorem statement, no proof, no proposal construction, no constants. That's not necessarily the authors' fault—it looks like a corruption in the pipeline—but it means my honest verdict is 'unverified.'\n\nOn the abstract alone, the soft spots are the usual ones for IS results. First, 'practical regimes of r/n' is not defined; if the constants in the bound depend on n in some hidden way, the 'complexity in r' headline could be misleading. Second, the proposal distribution is not described; importance sampling is only useful when the proposal overlaps the conditional distribution over mutational histories, and a poor proposal can have unbounded variance. The abstract gives no hint about how they construct it or what conditions it satisfies.\n\nThe positives: the upper/lower matching bounds are unusual and, if real, would be a solid technical contribution; the problem is well-motivated; the claim is specific enough to be falsified by reading the full text. I don't see any sign of circularity or fitting-to-data in the abstract.\n\nMy recommendation: if a clean, readable version of the paper is available, yes, send it to peer review. The claims are checkable and the topic is important enough to warrant referee time. But I wouldn't cite or rely on it until the full text is verified. For the reading group, it's a maybe: the abstract is worth discussing, but we can't do a real analysis on corrupted text.","headline":"A plausible importance-sampling estimator with matching finite-sample bounds, but the supplied full text is corrupt and unreadable, so the central claims are unverified.","tokens_in":23547,"tokens_out":2493,"would_cite":false,"duration_ms":28222,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An importance-sampling estimator approximates context-dependent sequence likelihoods with finite-sample error bounded above and below by matching orders, at a cost set by observed mutations rather than sequence length.","keywords":["importance sampling","sequence likelihood","context-dependent substitution models","site dependence","finite-sample error bounds","complexity in number of mutations","molecular evolution","randomized approximation"],"falsifier":"For a small context-dependent model where the exact likelihood can be enumerated, fix $r$ observed mutations and measure the estimator variance while increasing $n$; if the variance—or the sample count needed for fixed relative error—grows like $c^n$ rather than staying bounded in $n$, the central claim is wrong. Alternatively, take $r/n$ just inside the claimed practical regime and check whether the required sample size remains polynomial in $r$ or becomes exponential in $n$.","tokens_in":22669,"feed_emoji":"🧬","tokens_out":7380,"duration_ms":77897,"temperature":0.7,"pith_summary":"This paper claims that the marginal sequence likelihood under site-dependent (context-dependent) evolution models—where the substitution rate at one site depends on neighboring sites—can be approximated by an importance-sampling estimator whose finite-sample error is bounded above and below by matching orders. The payoff is a complexity statement: for two sequences of length $n$ with $r$ observed mutations, and for practical regimes of $r/n$, the sampler's cost grows in $r$, not exponentially in $n$. If correct, this opens likelihood-based inference for dependent-site models that exact computation cannot reach, and it provides a template for deriving model-specific complexity bounds. The authors demonstrate the template on a known dependent-site model from the phylogenetics literature.","feed_headline":"Mutation count, not genome length, sets sampling cost","feed_subtitle":"Matching upper and lower error bounds make context-dependent likelihoods computable when few sites differ.","key_machinery":"The load-bearing object is an importance-sampling estimator of the marginal sequence likelihood, built by proposing latent substitution histories between the observed sequences and reweighting them so that their average equals the target likelihood. The analysis turns on matching upper and lower bounds for the finite-sample approximation error, expressed in terms of the observed mutation count $r$ and the local context length of the model. This is what converts the estimator from a heuristic into a method with a stated complexity: the number of samples needed scales with $r$ rather than with $2^n$ or similar sequence-length factors in the sparse-mutation regime.","core_discovery":"The central claim is that computing the marginal probability of two observed sequences under a context-dependent substitution model can be done by randomized importance sampling with a finite-sample error that can be bounded from above and below in matching order. The estimator integrates over the unobserved mutational history connecting the two sequences; its error is controlled by the number of observed differences $r$ rather than by the sequence length $n$, provided the mutation count is not too large relative to $n$. In that sparse-mutation regime the sampler does not suffer the exponential-in-$n$ complexity that plagues exact and naive methods. The paper makes the bound concrete for a w","pith_inferences":["The paper leaves the boundary of the 'practical regime' of $r/n$ unspecified; mapping that threshold as a function of context length and substitution rate would turn the asymptotic guarantee into a usable acceptance criterion.","The same construction may extend to phylogenetic trees: if $r$ is reinterpreted as the total number of substitutions on all branches, the cost could scale with evolutionary divergence rather than alignment length, which would matter for ancestral-sequence inference on large trees.","A useful empirical check of the lower bound is to compare the estimator's variance against the predicted order for small models where the exact likelihood is computable; agreement would confirm the bound is tight, while large gaps would suggest the worst-case bound is pessimistic."],"forward_implications":["Likelihood evaluations become feasible for closely related long sequences under context-dependent models, enabling model comparison and parameter estimation that were previously out of reach.","For fixed $r$, increasing the sequence length $n$ should not increase the number of importance samples needed—so whole-genome comparisons with few differences are the natural operating range.","The matching lower bound indicates the error order is intrinsic to this sampling formulation; practical improvements should target the proposal distribution or variance reduction rather than sample count alone.","The general template lets practitioners derive their own complexity bound for a specific dependent-site model by computing the relevant model-dependent quantities, as done for the phylogenetics example."],"supporting_citations":[],"fun_headline_variants":["Sampling cost scales with mutations, not sequence length","Dependent-site likelihoods: error bounds scale with mutations","Importance sampling makes dependent-site models practical","Complexity in r, not n: sampling for site-dependent evolution"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The guarantees hold only when the importance-sampling proposal overlaps the target conditional distribution well enough to keep variance bounded, and only inside the paper's loosely defined practical regime of small $r/n$; outside either condition the complexity-in-$r$ claim has no force.","fun_headline_variants_meta":{"raw":{"variants":["Sampling cost scales with mutations, not sequence length","Dependent-site likelihoods: error bounds scale with mutations","Importance sampling makes dependent-site models practical","Complexity in r, not n: sampling for site-dependent evolution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00023,"raw_usage":{"total_tokens":1266,"prompt_tokens":635,"completion_tokens":631,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":379,"completion_tokens_details":{"reasoning_tokens":566}},"tokens_in":379,"tokens_out":631,"duration_ms":7245,"temperature":1.0,"reasoning_tokens":566,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:53:52.047290+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a small context-dependent model where the exact likelihood can be enumerated, fix $r$ observed mutations and measure the estimator variance while increasing $n$; if the variance—or the sample count needed for fixed relative error—grows like $c^n$ rather than staying bounded in $n$, the central claim is wrong. Alternatively, take $r/n$ just inside the claimed practical regime and check whether the required sample size remains polynomial in $r$ or becomes exponential in $n$.","supporting_citations":[],"review_version":1}