{"id":"9a679542-1389-4ffa-bbe0-0c211fafb013","arxiv_id":"2510.20500","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Using simulated evolution with known fitness, this study shows that MPL and tQLE both recover fitness order in many parameter regimes and mostly agree, with tQLE stronger under epistasis and MPL under drift.","lead":"This paper tests two statistical methods for inferring how natural selection acts on genes from time-stamped genome data, using simulated populations where the true fitness is known. It maps when these methods work and when they fail, helping researchers decide whether fitness can be recovered from real pathogen or ancient DNA time series.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fig. 5 and Appendix B report opposite MPL-vs-tQLE top-5% rank correlations; the 'either method' conclusion is not supported until this contradiction is resolved.","rationale":"I read the paper as a simulation benchmark intended to validate that two inference schemes recover fitness order in appropriate regimes. The load-bearing condition for the central claim is that the MPL-vs-tQLE comparison is internally consistent. It is not: the main-text Fig. 5 and Appendix B disagree on the direction of the advantage for top-5% rank at σ(f_i)=0.05 and 0.1. This is not a matter of differing evaluation criteria—both are Spearman rank correlations on top-5% sequences. The manuscript does not provide code, data, or the numeric values behind Fig. 5, so the contradiction cannot be resolved from the text. The acknowledged limitation of quadratic fitness landscapes is real but secondary; a resolution of the Fig. 5/Appendix B conflict is a prerequisite for any claim about real-data convenience. The reader's conditional verdict already captures the need for revision, so I do not change it; I agree partially, because my concern is the internal inconsistency rather than the quadratic-fitness assumption.","tokens_in":12903,"tokens_out":5892,"duration_ms":55109,"concrete_test":"Run the Table II parameter set (N=1000, L=25, T=30, r=0.5, μ=0.01, σ(f_ij)=0.002, σ(f_i)=0.005/0.05/0.1, 30 replicates) and compute the Spearman correlations plotted in Fig. 5 and Appendix A/B from the same simulated trajectories; report per-replicate and pooled values for all genotypes and for the top 5%. Release the code and data. If Appendix B's values reproduce (tQLE better at all σ), then Fig. 5/main text is wrong; if Fig. 5's values reproduce (MPL better at high σ), then Appendix B is wrong. Either outcome pinpoints which parameter regimes actually support the 'largely a matter of convenience' conclusion.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that fitness inference works with both MPL and tQLE and the choice is 'largely a matter of convenience'—rests on the comparison in Fig. 5. That figure is contradicted by the paper's own appendix. Fig. 8 (Appendix B) gives Spearman r for top-5% ranks: MPL/tQLE = 0.16/0.68 at σ(f_i)=0.005, 0.10/0.63 at 0.05, and 0.13/0.46 at 0.1; tQLE is better at every tested σ. The main text (Section III.B.1) says MPL 'becomes comparatively more robust' and the Fig. 5 caption says for high additive fitness 'MPL works consistently better,' and for intermediate σ that MPL outperforms tQLE for all sequences. At σ=0.1, Appendix A (full fitness) does show MPL better (0.91 vs 0.78), but Appendix B's top-5% rank shows the opposite; Fig. 5's dashed lines are the top-5% comparison, so the contradiction is direct. No code/data or numerical values for Fig. 5 are supplied, so the reader cannot decide which result is correct. If the appendix numbers stand, MPL never beats tQLE on top-5% rank in the tested range, which weakens the 'either method' conclusion. This internal inconsistency must be resolved before the comparative claim is used.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an in silico benchmark of two fitness-inference methods, the marginal path likelihood (MPL) method and the transient quasi-linkage equilibrium (tQLE) method, using FFPopSim simulations with known ground-truth fitness landscapes. The authors consider both additive-only and additive-plus-pairwise-epistatic fitness, evaluate recovery of individual fitness parameters and of genotype fitness order, and focus especially on the top 5% most-fit sequences. They conclude that both methods can infer fitness in appropriate parameter ranges, that they often agree with each other and with ground truth, and that for real data the choice between MPL and tQLE is 'largely a matter of convenience.'","tokens_in":13299,"tokens_out":9742,"duration_ms":80373,"significance":"If the reported results are correct, the paper provides a valuable validation of the authors' earlier empirical agreement between MPL and tQLE on SARS-CoV-2 data, showing that the agreement is not purely a shared error. The use of ground-truth simulations, multiple replicates for the epistatic case, and explicit rank-based evaluation are strengths. The paper honestly states its main limitations: no higher-order epistasis, well-mixed populations, uniform mutation and recombination rates. However, the internal inconsistencies between figures, the unstated MPL hyperparameter, and the narrow parameter coverage undermine confidence in the broad conclusion.","major_comments":[{"comment":"The top-5% rank correlations are contradictory. Fig. 5's caption states 'MPL works consistently better' at the highest σ(f_i)=0.1, but Fig. 8 reports Spearman r=0.13 for MPL and r=0.46 for tQLE, i.e., tQLE is substantially better. At σ(f_i)=0.05, Fig. 8 gives 0.63 vs 0.10, which is not 'marginally better' as claimed in Fig. 5. No numerical values are provided for Fig. 5's dashed lines, so the discrepancy cannot be resolved. This inconsistency affects the Discussion's conclusion that the choice of method is 'largely a matter of convenience'; if the appendix values are correct, MPL never beats tQLE on top-5% rank in the tested range.","section":"Section III.B.1 / Fig. 5 / Appendix B (Fig. 8)"},{"comment":"The MPL inference formula depends on the Gaussian-prior width γ, but the manuscript never states its value or the procedure used to choose it. Since the paper's core is a quantitative comparison of MPL and tQLE, the missing hyperparameter makes the MPL results irreproducible and the comparison not well-defined. Please report γ for every simulation, or state the criterion used to set it.","section":"Section II.B, Eq. (13)"},{"comment":"The conclusion that 'fitness inference is possible using both MPL and tQLE, and that it is largely a matter of convenience' is broader than the evidence. Simulations cover only one population size (N=1000), short genomes (L=25), 30 generations, strong recombination (r=0.5), and uniform mutation; Eq. (1) excludes higher-order epistasis and the model is well-mixed. These limitations are acknowledged, but the 'matter of convenience' claim does not follow from the presented results, especially because Fig. 8 indicates tQLE outperforms MPL on the practically important top-5% rank criterion across all tested σ(f_i). The conclusion should be restricted to the tested regimes and evaluation criteria, or supported by additional simulations.","section":"Section IV / Table I"},{"comment":"At σ(f_i)=0.05, Fig. 5's caption says 'MPL outperforms tQLE for all sequences', whereas Appendix A reports a higher correlation for tQLE (0.92 vs 0.87). If the former is a Spearman rank correlation and the latter a Pearson correlation, the text should say so explicitly; as written, the reader cannot tell whether the two statements are consistent.","section":"Section III.B.1 / Appendix A (Fig. 7)"}],"minor_comments":[{"comment":"The double sum in Eq. (1) should specify i<j for clarity.","section":"Section I / Eq. (1)"},{"comment":"The footnote says 'In Fig 3, Fig 4 and Fig 6 ... at non-zero epistatic fitness parameters F_ij', but Fig. 3 is the additive-only case with σ(f_ij)=0. Please correct.","section":"Table I footnote"},{"comment":"The term 'correlation' is used without specifying whether Pearson or Spearman. Given the paper's emphasis on rank-based evaluation, this should be stated explicitly.","section":"Appendix A / Fig. 7"},{"comment":"No code or data availability statement is provided. For a benchmark paper, releasing simulation scripts and processed results would substantially improve reproducibility.","section":"General"},{"comment":"The captions say 'Upper: μ=0.003, σ(f_i)=0.01, middle: μ=0.01, σ(f_i)=0.01, bottom: μ=0.01, σ(f_i)=0.05'. It would be helpful to indicate the same parameter ordering in the main text when discussing 'weaker' vs 'stronger' selection/mutation, to avoid ambiguity.","section":"Section III.A / Fig. 1-3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for q-bio.PE and reports an in silico benchmark that is potentially useful. The main concern is the unresolved internal inconsistency between Fig. 5 and Appendix B regarding top-5% rank correlations, which directly affects the central claim. The missing MPL hyperparameter value is also a serious reproducibility issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this paper gives a systematic in silico comparison of two fitness inference methods, MPL and tQLE, with known ground truth, and that is worth having. But the main conclusion—that the choice between methods is 'largely a matter of convenience'—is not supported by the paper's own numbers. The appendix and the main text disagree on which method recovers the top 5% fitness ranks, and the authors don't supply code or data to resolve it.\n\nWhat is new: prior work compared these methods on real SARS-CoV-2 data without ground truth. Here they simulate populations with FFPopSim, vary mutation rate, additive fitness variance, and include pairwise epistasis, then measure rank correlations against true fitness. That parameter sweep is a legitimate contribution. The paper is also honest about its main limitation: fitness is quadratic (no higher-order epistasis), and populations are well-mixed. Good.\n\nThe soft spots: Fig. 5 (main text) reported that MPL 'works consistently better' at high σ(f_i) and outperforms tQLE for all sequences at intermediate σ. Appendix B (Fig. 8) shows the opposite for the top 5%: tQLE beats MPL at every σ, with rank correlations like 0.46 vs 0.13 at σ=0.1. Both use 30 replicates and identical parameters, so this is a direct contradiction. The Discussion's 'either method' claim is load-bearing and it collapses if the appendix is right. Fig. 5 has no numerical values or error bars, and no code is provided—so the reader cannot tell which result is correct. Second, the claim that tQLE is 'dominated by MPL when genetic drift is important' is not tested: N=1000 throughout, so drift strength is never varied. Minor: the Gumbel fits in Fig. 4 are self-described as 'somewhat moot' because they fit minima, which is odd but not fatal.\n\nIf the appendix numbers stand, the appropriate conclusion is qualified: tQLE is better at ranking the top sequences in the tested range, MPL may be better for global ranks at high additive variance, and the convenience claim is unsupported. That is still a useful map.\n\nWho should read it: anyone applying MPL or tQLE to pathogen or experimental evolution data. It deserves a serious referee, but the authors should be asked to reconcile Fig. 5 with Appendix B and release the simulation code. I would not cite the comparative claim until that's resolved.\n\nRecommendation: send to peer review with a request for major revision.","headline":"A useful simulation benchmark whose central comparative claim is directly contradicted by its own appendix.","tokens_in":13741,"tokens_out":3052,"would_cite":true,"duration_ms":28309,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fitness inference from time-stratified genomes works with either of two methods across wide parameter ranges.","keywords":["Evolution","Fitness inference","Marginal path likelihood method","Transient quasi-linkage equilibrium method","in silico population genetics","epistasis","genotype fitness order","time-series genomic data"],"falsifier":"A forward simulation in which the ground-truth fitness landscape includes three-locus interaction terms (e.g., random Gaussian coefficients for f_ijk), keeping all other parameters in the recoverable range; if neither MPL nor tQLE recovers the true genotype fitness order, or if their predictions diverge, the paper's central 'convenience' claim collapses for biologically realistic landscapes.","tokens_in":12845,"feed_emoji":"🧬","tokens_out":4241,"duration_ms":37140,"temperature":0.7,"pith_summary":"The paper asks whether fitness parameters and genotype fitness order can be inferred from population-wide, time-stratified sequence data. To test this, the authors ran forward simulations of a well-mixed population under selection, mutation, recombination, and drift, with a known quadratic fitness function (additive plus pairwise epistasis). They compared two complementary inference approaches: marginal path likelihood (MPL), based on allele-frequency diffusion, and transient quasi-linkage equilibrium (tQLE), based on genome-wide statistical physics. Their central conclusion is that both methods recover the true fitness order of genotypes within broad parameter regimes, and that the two methods largely agree despite starting from different assumptions. This matters because it suggests that agreement between the methods on real data, such as SARS-CoV-2 genomes, reflects actual fitness signal rather than a shared bias.","feed_headline":"Fitness inference works with either of two methods, simulations show","feed_subtitle":"Ground-truth simulations map when MPL and tQLE recover fitness order, supporting their use on real genomic time series.","key_machinery":"The central object is the quadratic fitness model F(g) = Σ_i f_i s_i + Σ_{ij} f_ij s_i s_j, which limits the landscape to additive and pairwise epistatic effects and anchors both inference methods. MPL infers additive fitness coefficients from a diffusion approximation of allele-frequency trajectories, using a Fokker-Planck equation and a Gaussian-prior path likelihood (equation 13). tQLE assumes the evolving population is transiently described by a Gibbs-Boltzmann distribution over genomes, whose Ising parameters (single-site h and pair couplings J) are related to fitness parameters through quasi-linkage-equilibrium theory (equations 3–8). The pairing of these two different formalisms is wh","core_discovery":"On the paper's own terms, the discovery is that fitness inference is feasible with both MPL and tQLE, and that the choice between them is largely a matter of convenience on real data. Under additive fitness, both methods accurately infer individual selection coefficients and genotype fitness, with the main failure mode being weak selection relative to mutation. Under weak pairwise epistasis, tQLE is the better predictor of overall fitness when epistatic contributions dominate, while MPL catches up or surpasses it as additive variance grows; for the top 5% highest-fitness sequences, tQLE remains better across a range of parameters. The authors interpret the two methods' broad agreement on sim","pith_inferences":["A natural extension is to test whether the agreement between MPL and tQLE persists when the fitness landscape includes third- and higher-order epistasis; if not, the 'convenience' conclusion would only hold for the quadratic class.","The finding that tQLE outperforms MPL for top-ranked sequences even when global rank correlation favors MPL suggests that different evaluation metrics can lead to different 'best' methods; users should choose a metric tied to their ultimate goal.","The framework could be turned into a design tool: before collecting new time-series data, one could simulate a plausible fitness landscape to check whether the expected signal-to-noise ratio falls in the recoverable range.","The side result that epistatic fitness is not heritable under high recombination, while additive fitness is, could be tested experimentally in evolve-and-resequence experiments, linking the inference framework to measurable heritability."],"forward_implications":["If both methods recover fitness order, then for real datasets where MPL and tQLE agree—like the SARS-CoV-2 genomes analyzed earlier—the inferred fitness order is likely genuine, not a shared artifact.","A researcher can choose between MPL and tQLE based on convenience (data format, computational resources, whether epistatic coupling is needed) without losing accuracy, except when drift or strong pairwise epistasis dominates.","The simulation parameter maps give concrete guidance on when inference fails: weak selection relative to mutation, or strong drift not captured by tQLE, so users know when to distrust results.","The ability to rank the top 5% fittest genotypes from time-stratified data is especially promising for protein-engineering and pathogen-adaptation studies, where the most fit variants matter most."],"fun_headline_variants":["Sims map when fitness inference works from genomes","MPL vs tQLE: simulations show when each wins","Ground-truth sims test fitness inference methods","Two methods, one map: fitness inference limits"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The inference framework assumes fitness is a quadratic function of the genome with no higher-order epistasis; if real fitness landscapes contain substantial three-way or higher interactions, the ranking guarantees and the comparison between MPL and tQLE may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Sims map when fitness inference works from genomes","MPL vs tQLE: simulations show when each wins","Ground-truth sims test fitness inference methods","Two methods, one map: fitness inference limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1184,"prompt_tokens":620,"completion_tokens":564,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":364,"completion_tokens_details":{"reasoning_tokens":511}},"tokens_in":364,"tokens_out":564,"duration_ms":5438,"temperature":1.0,"reasoning_tokens":511,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T08:24:35.218365+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A forward simulation in which the ground-truth fitness landscape includes three-locus interaction terms (e.g., random Gaussian coefficients for f_ijk), keeping all other parameters in the recoverable range; if neither MPL nor tQLE recovers the true genotype fitness order, or if their predictions diverge, the paper's central 'convenience' claim collapses for biologically realistic landscapes.","supporting_citations":[],"review_version":1}