REVIEW 4 major objections 5 minor 50 references
Fitness inference tested by in silico population genetics
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Fitness inference from time-stratified genomes works with either of two methods across wide parameter ranges.
desk verdict A useful simulation benchmark whose central comparative claim is directly contradicted by its own appendix. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the quadratic fitness model F(g) = Σ_i f_i s_i + Σ_{ij} f_ij s_i s_j, which limits the landscape to additive and pairwise epistatic effects and anchors both inference methods. MPL infers additive fitness coefficients from a diffusion approximation of allele-frequency trajectories, using a Fokker-Planck equation and a Gaussian-prior path likelihood (equation 13). tQLE assumes the evolving population is transiently described by a Gibbs-Boltzmann distribution over genomes, whose Ising parameters (single-site h and pair couplings J) are related to fitness parameters through quasi-linkage-equilibrium theory (equations 3–8). The pairing of these two different formalisms is wh
What would settle it
A forward simulation in which the ground-truth fitness landscape includes three-locus interaction terms (e.g., random Gaussian coefficients for f_ijk), keeping all other parameters in the recoverable range; if neither MPL nor tQLE recovers the true genotype fitness order, or if their predictions diverge, the paper's central 'convenience' claim collapses for biologically realistic landscapes.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that fitness inference is feasible with both MPL and tQLE, and that the choice between them is largely a matter of convenience on real data. Under additive fitness, both methods accurately infer individual selection coefficients and genotype fitness, with the main failure mode being weak selection relative to mutation. Under weak pairwise epistasis, tQLE is the better predictor of overall fitness when epistatic contributions dominate, while MPL catches up or surpasses it as additive variance grows; for the top 5% highest-fitness sequences, tQLE remains better across a range of parameters. The authors interpret the two methods' broad agreement on sim
Load-bearing premise
The inference framework assumes fitness is a quadratic function of the genome with no higher-order epistasis; if real fitness landscapes contain substantial three-way or higher interactions, the ranking guarantees and the comparison between MPL and tQLE may not transfer.
Editorial extensions
If this is right
- If both methods recover fitness order, then for real datasets where MPL and tQLE agree—like the SARS-CoV-2 genomes analyzed earlier—the inferred fitness order is likely genuine, not a shared artifact.
- A researcher can choose between MPL and tQLE based on convenience (data format, computational resources, whether epistatic coupling is needed) without losing accuracy, except when drift or strong pairwise epistasis dominates.
- The simulation parameter maps give concrete guidance on when inference fails: weak selection relative to mutation, or strong drift not captured by tQLE, so users know when to distrust results.
- The ability to rank the top 5% fittest genotypes from time-stratified data is especially promising for protein-engineering and pathogen-adaptation studies, where the most fit variants matter most.
Reading between the lines
- A natural extension is to test whether the agreement between MPL and tQLE persists when the fitness landscape includes third- and higher-order epistasis; if not, the 'convenience' conclusion would only hold for the quadratic class.
- The finding that tQLE outperforms MPL for top-ranked sequences even when global rank correlation favors MPL suggests that different evaluation metrics can lead to different 'best' methods; users should choose a metric tied to their ultimate goal.
- The framework could be turned into a design tool: before collecting new time-series data, one could simulate a plausible fitness landscape to check whether the expected signal-to-noise ratio falls in the recoverable range.
- The side result that epistatic fitness is not heritable under high recombination, while additive fitness is, could be tested experimentally in evolve-and-resequence experiments, linking the inference framework to measurable heritability.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an in silico benchmark of two fitness-inference methods, the marginal path likelihood (MPL) method and the transient quasi-linkage equilibrium (tQLE) method, using FFPopSim simulations with known ground-truth fitness landscapes. The authors consider both additive-only and additive-plus-pairwise-epistatic fitness, evaluate recovery of individual fitness parameters and of genotype fitness order, and focus especially on the top 5% most-fit sequences. They conclude that both methods can infer fitness in appropriate parameter ranges, that they often agree with each other and with ground truth, and that for real data the choice between MPL and tQLE is 'largely a matter of convenience.'
Significance. If the reported results are correct, the paper provides a valuable validation of the authors' earlier empirical agreement between MPL and tQLE on SARS-CoV-2 data, showing that the agreement is not purely a shared error. The use of ground-truth simulations, multiple replicates for the epistatic case, and explicit rank-based evaluation are strengths. The paper honestly states its main limitations: no higher-order epistasis, well-mixed populations, uniform mutation and recombination rates. However, the internal inconsistencies between figures, the unstated MPL hyperparameter, and the narrow parameter coverage undermine confidence in the broad conclusion.
major comments (4)
- [Section III.B.1 / Fig. 5 / Appendix B (Fig. 8)] The top-5% rank correlations are contradictory. Fig. 5's caption states 'MPL works consistently better' at the highest σ(f_i)=0.1, but Fig. 8 reports Spearman r=0.13 for MPL and r=0.46 for tQLE, i.e., tQLE is substantially better. At σ(f_i)=0.05, Fig. 8 gives 0.63 vs 0.10, which is not 'marginally better' as claimed in Fig. 5. No numerical values are provided for Fig. 5's dashed lines, so the discrepancy cannot be resolved. This inconsistency affects the Discussion's conclusion that the choice of method is 'largely a matter of convenience'; if the appendix values are correct, MPL never beats tQLE on top-5% rank in the tested range.
- [Section II.B, Eq. (13)] The MPL inference formula depends on the Gaussian-prior width γ, but the manuscript never states its value or the procedure used to choose it. Since the paper's core is a quantitative comparison of MPL and tQLE, the missing hyperparameter makes the MPL results irreproducible and the comparison not well-defined. Please report γ for every simulation, or state the criterion used to set it.
- [Section IV / Table I] The conclusion that 'fitness inference is possible using both MPL and tQLE, and that it is largely a matter of convenience' is broader than the evidence. Simulations cover only one population size (N=1000), short genomes (L=25), 30 generations, strong recombination (r=0.5), and uniform mutation; Eq. (1) excludes higher-order epistasis and the model is well-mixed. These limitations are acknowledged, but the 'matter of convenience' claim does not follow from the presented results, especially because Fig. 8 indicates tQLE outperforms MPL on the practically important top-5% rank criterion across all tested σ(f_i). The conclusion should be restricted to the tested regimes and evaluation criteria, or supported by additional simulations.
- [Section III.B.1 / Appendix A (Fig. 7)] At σ(f_i)=0.05, Fig. 5's caption says 'MPL outperforms tQLE for all sequences', whereas Appendix A reports a higher correlation for tQLE (0.92 vs 0.87). If the former is a Spearman rank correlation and the latter a Pearson correlation, the text should say so explicitly; as written, the reader cannot tell whether the two statements are consistent.
minor comments (5)
- [Section I / Eq. (1)] The double sum in Eq. (1) should specify i<j for clarity.
- [Table I footnote] The footnote says 'In Fig 3, Fig 4 and Fig 6 ... at non-zero epistatic fitness parameters F_ij', but Fig. 3 is the additive-only case with σ(f_ij)=0. Please correct.
- [Appendix A / Fig. 7] The term 'correlation' is used without specifying whether Pearson or Spearman. Given the paper's emphasis on rank-based evaluation, this should be stated explicitly.
- [General] No code or data availability statement is provided. For a benchmark paper, releasing simulation scripts and processed results would substantially improve reproducibility.
- [Section III.A / Fig. 1-3] The captions say 'Upper: μ=0.003, σ(f_i)=0.01, middle: μ=0.01, σ(f_i)=0.01, bottom: μ=0.01, σ(f_i)=0.05'. It would be helpful to indicate the same parameter ordering in the main text when discussing 'weaker' vs 'stronger' selection/mutation, to avoid ambiguity.
Circularity Check
No significant circularity: MPL and tQLE are validated against FFPopSim ground-truth simulations, not fitted to the target quantities.
full rationale
This is a benchmark study, not a circular derivation. The true fitness parameters f_i and f_ij are drawn randomly and used only to generate simulated genotype time series in FFPopSim; the inference formulas (Eqs. 8 and 13) are applied to the simulated allele-frequency data and produce fitness estimates that are then compared to the known ground truth. No inference parameter is fitted to the target fitness values, and no predicted quantity is defined in terms of the ground truth. The MPL and tQLE methods are imported from prior papers, some with overlapping authors, but those citations supply the inference formulas and their theory, not the validation; the validation here is external to the self-citations because it uses an independent simulator with known fitness. Self-citation is therefore not load-bearing for the paper's central claim. The only notable issue is an internal inconsistency between the main-text description of Fig. 5 and the numerical Spearman correlations reported in Appendix B for top-5% rank recovery; that is a correctness/consistency concern rather than a circularity concern. The derivation chain is self-contained with respect to circularity, so the appropriate score is 0.
Assumptions & free parameters
free parameters (1)
- gamma (MPL Gaussian-prior width)
assumptions (6)
- domain assumption Fitness is quadratic with pairwise epistasis only (Eq. 1)
- domain assumption Biallelic loci and well-mixed population with uniform mutation and recombination
- standard math Kimura diffusion approximation is adequate for MPL
- domain assumption Transient quasi-linkage equilibrium (Gibbs-Boltzmann form) holds for tQLE
- domain assumption FFPopSim faithfully simulates selection/mutation/recombination/drift
- domain assumption Selection and mutation are comparable to or stronger than drift (N=1000)
Cite this review
Pith. "Pith review of Fitness inference tested by in silico population genetics." pith.science (2026). https://pith.science/paper/Z66PSJ7U
@misc{pith2026251020500,
author = {Pith},
title = {Pith review of: Fitness inference tested by in silico population genetics},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z66PSJ7U}},
note = {Machine review of arXiv:2510.20500}
}
read the original abstract
We consider populations evolving according to natural selection, mutation, and recombination, and assume that the genomes of all or a representative selection of individuals are known. We pose the problem if it is possible to infer fitness parameters and genotype fitness order from such data. We tested this hypothesis in simulated populations. We delineate parameter ranges where this is possible and other ranges where it is not.Our work provides a framework for determining when fitness inference is feasible from population-wide, whole-genome, time-stratified data and highlights settings where it is not. We give a brief survey of biological model organisms and human pathogens that fit into this framework.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Recovery of genotype-level fitness We used Spearman correlations to quantify the overall correspondence between inferred and true fitness ranks across all genotypes and, separately, within the elite top 5% of true fitness values (Fig. 5). At lowσ(f i), tQLE yields higher global rank correlations than MPL. Asσ(f i) increases, MPL becomes comparatively more...
-
[2]
Figure 6 shows that both approaches obtain a good linear correlation between the true and inferred additive contributions to fitnessf i
Recovery of individual fitness parameters In addition to genotype-level fitness estimates, we also assessed the ability of tQLE and MPL to recover underlying additive and epistatic contributions to fitness from simulated evolutionary trajectories. Figure 6 shows that both approaches obtain a good linear correlation between the true and inferred additive c...
2020
-
[3]
Mathieson, I
I. Mathieson, I. Lazaridis, N. Rohland, S. Mallick, N. Patterson, S. A. Roodenberg, E. Harney, K. Stewardson, D. Fernandes, M. Novak,et al., Nature528, 499 (2015)
2015
-
[4]
W. J. Ewens,Mathematical population genetics: theoretical introduction, Vol. 27 (Springer, 2004)
2004
-
[5]
J. P. Bollback, T. L. York, and R. Nielsen, Genetics179, 497 (2008)
2008
-
[6]
Lacerda and C
M. Lacerda and C. Seoighe, Genetics198, 1237 (2014)
2014
-
[7]
Malaspinas, O
A.-S. Malaspinas, O. Malaspinas, S. N. Evans, and M. Slatkin, Genetics192, 599 (2012)
2012
-
[8]
Mathieson and G
I. Mathieson and G. McVean, Genetics193, 973 (2013)
2013
Show all 50 references
-
[9]
A. F. Feder, S. Kryazhimskiy, and J. B. Plotkin, Genetics196, 509 (2014)
2014
-
[10]
Steinr¨ ucken, A
M. Steinr¨ ucken, A. Bhaskar, and Y. S. Song, The annals of applied statistics8, 2203 (2014)
2014
-
[11]
Ferrer-Admetlla, C
A. Ferrer-Admetlla, C. Leuenberger, J. D. Jensen, and D. Wegmann, Genetics203, 831 (2016)
2016
-
[12]
J. G. Schraiber, S. N. Evans, and M. Slatkin, Genetics203, 493 (2016)
2016
-
[13]
Paris, B
C. Paris, B. Servin, and S. Boitard, G3: Genes, Genomes, Genetics9, 4073 (2019)
2019
-
[14]
J. M. Smith and J. Haigh, Genetics Research23, 23 (1974)
1974
-
[15]
H. J. Muller, The American Naturalist66, 118 (1932)
1932
-
[16]
P. J. Gerrish and R. E. Lenski, Genetica102, 127 (1998)
1998
-
[17]
C. J. Illingworth and V. Mustonen, Genetics189, 989 (2011)
2011
-
[18]
C. J. Illingworth and V. Mustonen, Bioinformatics28, 831 (2012)
2012
-
[19]
Foll, Y.-P
M. Foll, Y.-P. Poh, N. Renzette, A. Ferrer-Admetlla, C. Bank, H. Shim, A.-S. Malaspinas, G. Ewing, P. Liu, D. Wegmann, et al., PLoS genetics10, e1004185 (2014)
2014
-
[20]
Terhorst, C
J. Terhorst, C. Schl¨ otterer, and Y. S. Song, PLoS genetics11, e1005069 (2015)
2015
-
[21]
Tataru, M
P. Tataru, M. Mollion, S. Gl´ emin, and T. Bataillon, Genetics207, 1103 (2017)
2017
-
[22]
M. S. Sohail, R. H. Y. Louie, M. R. McKay, and J. P. Barton, Nature Biotechnology39, 472 (2021)
2021
-
[23]
Zeng, C.-L
H.-L. Zeng, C.-L. Yang, B. Jing, J. Barton, and E. Aurell, Physical Biology22, 016003 (2024)
2024
-
[24]
M. S. Sohail, R. H. Louie, Z. Hong, J. P. Barton, and M. R. McKay, Molecular biology and evolution39, msac199 (2022)
2022
-
[25]
K. S. Shimagaki and J. P. Barton, Genetics230, iyaf118 (2025)
2025
-
[26]
Shu and J
Y. Shu and J. McCauley, Eurosurveillance22(2017)
2017
-
[27]
Weigt, R
M. Weigt, R. A. White, H. Szurmant, J. A. Hoch, and T. Hwa, Proceedings of the National Academy of Sciences106, 67 (2009), https://www.pnas.org/content/106/1/67.full.pdf
2009
-
[28]
Cocco, C
S. Cocco, C. Feinauer, M. Figliuzzi, R. Monasson, and M. Weigt, Reports on Progress in Physics81, 10.1088/1361- 6633/aa9965 (2018)
2018 doi
-
[29]
R. A. Fisher,The Genetical Theory of Natural Selection(Clarendon, 1930)
1930
-
[30]
R. A. Blythe and A. J. McKane, J. Stat. Mech.: Theory Exp.2007(07), P07018
2007
-
[31]
Kimura, Evolution10, 278 (1956)
M. Kimura, Evolution10, 278 (1956)
1956
-
[32]
Kimura, J
M. Kimura, J. Appl. Probab.1, 177–232 (1964)
1964
-
[33]
Pathogen detection BETA, https://www.ncbi.nlm.nih.gov/pathogens/ (2025)
2025
-
[34]
P¨ a¨ abo, H
S. P¨ a¨ abo, H. Poinar, D. Serre, V. Jaenicke-Despr´ es, J. Hebler, N. Rohland, M. Kuch, J. Krause, L. Vigilant, and M. Hofre- iter, Annu. Rev. Genet.38, 645 (2004)
2004
-
[35]
Orlando, R
L. Orlando, R. Allaby, P. Skoglund, C. Der Sarkissian, P. W. Stockhammer, M. C. ´Avila-Arcos, Q. Fu, J. Krause, E. Willerslev, A. C. Stone,et al., Nature reviews methods primers1, 14 (2021)
2021
-
[36]
K. H. Kjær, M. Winther Pedersen, B. De Sanctis, B. De Cahsan, T. S. Korneliussen, C. S. Michelsen, K. K. Sand, S. Jelavi´ c, A. H. Ruter, A. M. Schmidt,et al., Nature612, 283 (2022)
2022
-
[37]
R. E. Lenski, M. R. Rose, S. C. Simpson, and S. C. Tadler, The American Naturalist138, 1315 (1991)
1991
-
[38]
J. A. Ascensao and M. M. Desai, Nature Reviews Genetics , 1 (2025)
2025
-
[39]
Zanini and R
F. Zanini and R. A. Neher, Bioinformatics28, 3332 (2012)
2012
-
[40]
R. A. Neher and B. I. Shraiman, Rev. Mod. Phys.83, 1283 (2011)
2011
-
[41]
Zeng and E
H.-L. Zeng and E. Aurell, Phys. Rev. E101, 052409 (2020)
2020
-
[43]
H.-L. Zeng, Y. Liu, V. Dichio, and E. Aurell, Phys. Rev. E106, 044409 (2022)
2022
-
[44]
R. A. Neher, M. Vucelja, M. Mezard, and B. I. Shraiman, Journal of Statistical Mechanics: Theory and Experiment2013, P01008 (2013)
2013
-
[45]
C. Sire, S. N. Majumdar, and D. S. Dean, Journal of Statistical Mechanics: Theory and Experiment2006, L07001 (2006)
2006
-
[46]
Krug and K
J. Krug and K. Jain, Physica A: Statistical Mechanics and its Applications358, 1 (2005), condensed Matter and Statistical Physics
2005
-
[47]
Jain, Phys
K. Jain, Phys. Rev. E76, 031922 (2007)
2007
-
[48]
P¨ a¨ abo,Neanderthal Man: In Search of Lost Genomes(Basic Civitas Books, 2014)
S. P¨ a¨ abo,Neanderthal Man: In Search of Lost Genomes(Basic Civitas Books, 2014)
2014
-
[49]
Reich,Ancient DNA and the new science of the human past(Oxford University Press, 2019)
D. Reich,Ancient DNA and the new science of the human past(Oxford University Press, 2019)
2019
-
[50]
Kimura, Genetics52, 875 (1965)
M. Kimura, Genetics52, 875 (1965)
1965
-
[51]
H.-L. Zeng, E. Mauri, V. Dichio, S. Cocco, R. Monasson, and E. Aurell, Journal of Statistical Mechanics: Theory and Experiment2021, 083501 (2021)
2021
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.