REVIEW 3 major objections 5 minor 25 references
Exploring the Multifractal Behavior of the Human Genome T2T-CHM13v2.0: Graphical Representations and Cytogenetics
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The complete human genome has a multifractal structure that a 12-base Markov chain reproduces to within about 2 percent error.
desk verdict Worth a look for the descriptive T2T CGR images, but the quantitative multifractal spectra and 2% MC fit are not established: they compute the partition sum at one box size and never take the limit in Eq. (5). read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the multifractal spectrum f($\alpha$), obtained from a chaos game representation (CGR) of a genomic sequence: each base is mapped to a corner of the unit square and iterated as an iterated function system, producing a point set whose box-counting probabilities p_l lead to the partition function tau(q), then to the generalized dimension D_q and the spectrum via a Legendre transform. The box-counting coverage at resolutions $4^{12}$ (whole assembly) and $4^{10}$ (individual chromosomes) is what carries the analysis. The Markov chain representation uses transition probabilities P_{XY} of order n-1, with n=2,3,6,12, to generate surrogate sequences; the binary genomic representation encodes bases as two-bit vectors and decodes losslessly. The comparison between these representations is performed through their multifractal spectra.
What would settle it
Compute the multifractal spectrum of the same assembly at a higher resolution, such as $4^{14}$ boxes, or on bootstrap-subsampled halves of the genome, and check whether chromosome 9 and Y retain the widest spectra and whether a 12-base Markov chain still fits to within a few percent; if the spectra change substantially, the claimed common fractal support is an artifact of the chosen box count.
Extended reading notes
Core claim
The central claim is that the complete human genome has a well-defined multifractal structure that can be recovered, almost exactly, by a low-order Markov chain. Using chaos game representation and box-counting with $4^{12}$ boxes for the full assembly and $4^{10}$ boxes per chromosome, the paper derives multifractal spectra f($\alpha$) for each sequence. The spectra show a common geometric support across chromosomes but distinct probability distributions; chromosomes 9 and Y deviate most, with the widest singularity ranges, and the Y chromosome's spectrum maximum is slightly shifted. A Markov chain with dodecanucleotide transition probabilities, run to the full assembly length, reproduces the assembly's multifractal spectrum with a mean error near 2 percent, whereas a binary genomic representation matches better in high-frequency, non-coding regions but loses precision in low-incidence coding zones.
Load-bearing premise
The entire analysis assumes that box-counting the finite set of CGR points at one fixed resolution ($4^{12}$ for the assembly, $4^{10}$ for chromosomes) gives a stable, representative multifractal spectrum, with no convergence check, bootstrap uncertainty, or comparison to synthetic null sequences to confirm the spectra are not artifacts of sequence length.
Editorial extensions
If this is right
- If the result holds, the human genome's sequence organization can be summarized by a compact set of spectra and a 12-base Markov chain, enabling fast simulation of genome-like sequences for comparative or modeling purposes.
- The observed separation between coding and non-coding regions, and between CpG islands and the rest, within the CGR suggests that multifractal spectra could serve as a quantitative marker for annotating functional genomic elements.
- The characteristic banding patterns seen in the one-dimensional CGR plots align with cytogenetic banding, which could link sequence-level fractal measures to chromosome structure and, potentially, to chromosomal aberrations.
- Since mitochondrial DNA shows a distinct multifractal structure resembling a Sierpinski triangle, separate treatment of organelle genomes in multifractal analyses is warranted.
- The finding that chromosomes 9 and Y have the widest singularity spectra invites targeted study of these chromosomes for sequence features that generate such heterogeneity.
Reading between the lines
- A testable extension is to check whether the 2 percent Markov-chain fit survives at higher box resolutions (e.g., 4^14) or when the chain is trained on one chromosome and tested on another; the paper's claim that 12-base chains are optimal would be strengthened or weakened accordingly.
- The paper's implicit claim that the multifractal spectrum captures 'function' of genomic regions could be probed by correlating local spectrum widths with gene density, recombination rates, or histone modification maps, which the authors do not do.
- The BGR method's loss of precision in coding zones under high compression suggests a possible trade-off between lossless compressibility and preservation of low-frequency genomic features; this could be explored as a general principle for genomic encodings.
- Since the spectrum depends on the chosen box sizes and sequence length, a natural next step is to establish whether the 4^12/4^10 choice is stable across bootstrap resamples of the genome; the paper does not provide such a convergence analysis.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper applies Chaos Game Representation (CGR) to the complete T2T-CHM13v2.0 human genome assembly, individual chromosomes, and mitochondrial DNA. Multifractal spectra are computed via box-counting partitions at a single resolution (4^12 boxes for the assembly, 4^10 for chromosomes), and the spectra are compared with those of Markov-chain-generated sequences and a proposed binary representation (BGR). The reported findings are that the fractal support is consistent across chromosomes, chromosomes 9 and Y show the widest singularity spectra, CGR densities approximately separate coding and non-coding regions as well as CpG islands, one-dimensional CGR projections reveal cytogenetic band-like patterns, and a dodecanucleotide Markov chain reproduces the assembly spectrum with an average error of about 2%.
Significance. The use of the complete T2T assembly is a strength, and the qualitative observations (e.g., CG dinucleotide depletion, chromosome-specific distributions) are consistent with prior literature on genomic CGRs. The BGR construction is lossless and the MC comparison is a reasonable modeling exercise. However, the central quantitative claims depend on a single-resolution box-counting implementation of Eq. (5), which does not test the required scaling limit, and on an undefined 2% error metric. Because these issues bear directly on the main conclusions, the significance of the paper is conditional on their resolution.
major comments (3)
- [Section II and Eq. (5)] Section II states that subdivisions of 4^12 boxes (assembly) and 4^10 boxes (chromosomes) were used, and τ(q) was derived from Eq. (5), which defines τ(q) as a limit r→0. No multi-resolution scaling analysis is reported: no linear regression of log Σ p_l^q against log r, no scaling range, and no test that the spectrum is stable under changes of grid size. The Legendre transform of a single-resolution partition sum is not a multifractal spectrum in the sense of Eq. (5); its shape changes with resolution, so the claims about chromosomes 9 and Y (Fig. 13) and the MC agreement (Fig. 19) may be artifacts of the chosen resolutions. Please provide a genuine scaling analysis with multiple box sizes and report the fitted exponents.
- [Section I C / Section II (empty boxes)] Eq. (5) uses p_l^q for q from -30 to 30. For negative q, boxes with p_l = 0 produce divergent contributions, and with finite point counts and 4^12 boxes many boxes are necessarily empty. The manuscript never states how empty boxes were treated (pseudocount, occupied-box restriction, or some other rule). This is not a technical footnote: the low-q regime is emphasized for the chromosome 9/Y comparisons and for the MC/BGR fits, and any zero-count rule changes the low-q tails. The authors must specify the rule and test the sensitivity of the spectra to it.
- [Section III B (MC fit and error metric)] The Markov chain parameters are estimated from the same T2T-CHM13v2.0 assembly using Eq. (1) and then the MC-generated CGR is compared against that same assembly's spectrum. This is a self-consistency check, not an independent prediction, so the 'fit' measures how well a Markov model captures the genome's own statistics. In addition, the 'average error of approximately 2%' is not defined: no formula for the error, no confidence interval, and no reported variation over MC realizations or subsamples is given. The central quantitative claim of the paper therefore rests on an undefined metric. Please define the error, report its distribution, and ideally compare against a held-out or synthetic control.
minor comments (5)
- [Section I B 3] The text around Eq. (3) contains 'n+1 = ⌊# de bases en la GS / M⌋ + 1' with Spanish phrase 'de bases en la GS' inside an otherwise English sentence; this should be translated and typeset properly.
- [Abstract and Section III B] The acronym for the binary representation is given as 'RGB' in the abstract and as 'BGR' in the body; please standardize to a single abbreviation.
- [Section I C] The text refers to 'Monte Carlo (MC)' in the sentence 'the results obtained from Monte Carlo (MC) and BGR methods,' whereas MC is defined earlier as Markov Chain; this creates ambiguity and should be corrected.
- [References [12] and [17]] References [12] and [17] both cite Barnsley's 'Fractals Everywhere'; please consolidate to a single reference.
- [Section III A and Figure 10] The transformation described as 'x + ⌊y(46 − 1)⌋' and the phrase '46 rows' should read '4^6 rows' (i.e., 4096 rows), since the CGR is divided into 4^6 rows; the current notation is ambiguous.
Circularity Check
Markov-chain 'fit' to the assembly's multifractal spectrum is self-definitional: for n=12 the MC is parameterized by the same 12-mer frequencies that the 4^12 CGR box-counting measures.
-
self definitional
[Section III.B (MC and BGR Methods), with Eq. (1) and Section II Methodology]
"Precision significantly improves with longer sequences, reaching near-perfect fit with 12-base chains, suggesting that the optimal MC sequence length should equal the number of CGR partitions."
For n=12, Eq. (1) estimates P_XY from the counts of 12-mers in the T2T genome. The CGR is partitioned into 4^12 boxes (Section II), and each box corresponds to a distinct 12-base suffix; hence the p_l in Eq. (5) are the empirical 12-mer frequencies. The MC with n=12 is the maximum-likelihood order-11 Markov model of exactly those 12-mer frequencies, so an MC-generated sequence of equal length has, in expectation, the same box counts at 4^12 resolution. The multifractal spectra (Eqs. (5)-(7)) evaluated at this single resolution are therefore equal by construction up to sampling fluctuations. Reporting a 2% 'fit' is an in-sample consistency statement, not an independent confirmation that the MC simulates the genome's multifractal structure.
full rationale
The identified circularity is confined to the Markov-chain comparison: the n=12 MC is estimated from the same 12-mer frequency histogram that the 4^12 CGR box-counting spectrum is computed from, making the reported 2% agreement a self-consistency check rather than an independent test. The chromosome-level multifractal observations (chromosomes 9 and Y, cytogenetic bands) and the BGR comparison are empirical findings about the same dataset and do not reduce to their inputs by construction. Self-citations (refs. [14], [15]) are not load-bearing. A separate methodological concern—the use of a single box resolution rather than the r→0 limit in Eq. (5)—undermines the validity of the multifractal quantities, but that is a correctness risk, not a circularity. Because the MC 'prediction' is a central advertised result but the paper also contains independent descriptive content, a score of 6 reflects partial circularity rather than total collapse.
Assumptions & free parameters
assumptions (3)
- domain assumption The T2T-CHM13v2.0 assembly is a complete and accurate representation of the human genome, including the Y chromosome.
- domain assumption The finite CGR point set is long enough that box-counting at resolutions of 4^12 and 4^10 approximates the multifractal spectrum of the underlying genome measure.
- standard math The standard multifractal formalism (partition function, Renyi dimensions, Legendre transform) applies to the empirical measures defined by CGR point densities.
Cite this review
Pith. "Pith review of Exploring the Multifractal Behavior of the Human Genome T2T-CHM13v2.0: Graphical Representations and Cytogenetics." pith.science (2026). https://pith.science/paper/X5TORWTA
@misc{pith2026241216705,
author = {Pith},
title = {Pith review of: Exploring the Multifractal Behavior of the Human Genome T2T-CHM13v2.0: Graphical Representations and Cytogenetics},
year = {2026},
howpublished = {\url{https://pith.science/paper/X5TORWTA}},
note = {Machine review of arXiv:2412.16705}
}
read the original abstract
In this work, we applied the Chaos Game Representation (CGR) to the complete human genomic sequence T2T-CHM13v2.0, analyzing the entire chromosome assembly and each chromosome separately, including mitochondrial DNA. Multifractal spectra were determined using two types of box-counting coverage, revealing slight variations across most chromosomes. While the geometric support remained consistent, distinct distributions were observed for each chromosome. Chromosomes 9 and Y exhibited the greatest differences in singularity (H\"older exponent), with minor variations in their fractal support. The CGR distributions generally demonstrated an approximate separation between coding and non-coding sections, as well as CpG or GpC islands. A base-by-base analysis of the fractal support of the CGR uncovered characteristic structural bands in chromosome sequences, which align with patterns identified in cytogenetic studies. Using the complete assembly as a reference, we compared two alternative representations: the Binary Genomic Representation (RGB) and the Markov Chain (MC) representation. Both methods tended toward the same fractal support but displayed differing distributions based on the assigned length parameter. Multifractal analysis highlighted quantitative differences between these representations: RGB aligned more closely with high-frequency components, while MC showed better correspondence with low frequencies. The optimal fit was achieved using MC for twelve-base chains, yielding an average percentage error of 2% relative to the full genomic assembly.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
In our case, the method involves Figure 5: Example of point distribution in the CGR
Chaos Game Representation The game of chaos is a mathematical visualization method used to show the principles of chaos theory [13– 15] and fractals [16, 17]. In our case, the method involves Figure 5: Example of point distribution in the CGR. In region P23, the population is zero, which is reflected throughout the representation due to the absence of seg...
-
[2]
Markov Chains Here, we construct Markov chains by assigning proba- bilities of obtaining any of the bases conditioned on the previous state. These previous states are defined as X, which are short genomic sequences of fixed length n − 1 for n = 1, 2, 3, . . ., and the final state Y ∈ {A, C, G, T}. For example, for n = 2, XY is a dinucleotide, meaning a se...
-
[3]
Binary Genomic Representation Inspired by the ideas proposed by [18], we propose a new method for compressing and representing GSs, consisting of assigning two-dimensional binary vectors to each of the bases: A = 0 0 , C = 0 1 , G = 1 1 , T = 1 0 . (2) In this way, each point on the plane is represented in binary form, and the points in the BGR are define...
work page 2021
-
[4]
P. J. Deschavanne, A. Giron, J. Vilain, G. Fagot, and B. Fertil, “Genomic signature: characterization and clas- sification of species assessed by chaos game represen- tation of sequences.,” Molecular Biology and Evolution , vol. 16, pp. 1391–1399, 10 1999
work page 1999
-
[5]
National Human Genome Research Institute
NIH, “National Human Genome Research Institute.” https://www.genome.gov/, 2024
work page 2024
-
[6]
T2T, “Genome assembly T2T-CHM13v2.0.” https://www.ncbi.nlm.nih.gov/datasets/genome/ 15 GCF_009914755.1/, 2022
work page 2022
-
[7]
Chaos game representation of gene struc- ture,
H. Jeffrey, “Chaos game representation of gene struc- ture,” Nucleic Acids Research , vol. 18, pp. 2163–2170, 04 1990
work page 1990
-
[8]
G. Albrecht-Buehler, “Fractal genome sequences,” Gene, vol. 498, no. 1, pp. 20–27, 2012
work page 2012
Show all 25 references
-
[9]
Fractal properties of the human genome,
S. Garte, “Fractal properties of the human genome,” Journal of Theoretical Biology , vol. 230, no. 2, pp. 251– 260, 2004
2004
-
[10]
Characterizing fractal genetic variation in the hu- man genome from the hapmap project,
A. Borri, A. Cerasa, P. Tonin, L. Citrigno, and C. Por- caro, “Characterizing fractal genetic variation in the hu- man genome from the hapmap project,” International Journal of Neural Systems , vol. 32, no. 06, p. 2250028,
-
[11]
The human genome: a multifractal analysis,
P. A. Moreno, P. E. V´ elez, E. Mart ´ ınez, L. E. Garreta, N. D ´ ıaz, S. Amador, I. Tischer, J. M. Guti´ errez, A. K. Naik, F. Tobar, and F. Garc ´ ıa, “The human genome: a multifractal analysis,” BMC Genomics, vol. 12, pp. 1–17, 2011
2011
-
[12]
M. F. Barnsley, Fractals Everywhere. USA: Dover Pub- lications, Inc., 2012
2012
-
[13]
On the fractal geometry of dna by the binary image analysis,
C. Cattani and G. Pierro, “On the fractal geometry of dna by the binary image analysis,” Bulletin of Mathe- matical Biology, vol. 75, pp. 1544–1570, 2013
2013
-
[14]
Measure represen- tation and multifractal analysis of complete genomes,
Z.-G. Yu, V. Anh, and K.-S. Lau, “Measure represen- tation and multifractal analysis of complete genomes,” Physical Review E , vol. 64, no. 3, p. 031903, 2001
2001
-
[15]
Strachan and A
T. Strachan and A. Read, Human Molecular Genetics . CRC Press, 2018
2018
-
[16]
Feder, Fractals
J. Feder, Fractals. Springer Science & Business Media, 2013
2013
-
[17]
S. H. Strogatz, Nonlinear dynamics and chaos: with ap- plications to physics, biology, chemistry, and engineering. CRC press, 2018
2018
-
[18]
Classical harmonic three-body system: an experimental electronic realization,
A. Escobar-Ruiz, M. Quiroz-Juarez, J. Del Rio-Correa, and N. Aquino, “Classical harmonic three-body system: an experimental electronic realization,” Scientific Re- ports, vol. 12, no. 1, p. 13346, 2022
2022
-
[19]
Experimental realization of the classical dicke model,
M. A. Quiroz-Ju´ arez, J. Ch´ avez-Carlos, J. L. Arag´ on, J. G. Hirsch, and R. d. J. Le´ on-Montiel, “Experimental realization of the classical dicke model,” Physical Review Research, vol. 2, no. 3, p. 033169, 2020
2020
-
[20]
Harte, Multifractals: Theory and Applications
D. Harte, Multifractals: Theory and Applications . CRC Press, 2001
2001
-
[21]
M. F. Barnsley, Fractals everywhere. Academic press, 2014
2014
-
[22]
Encoding and decoding dna sequences by in- teger chaos game representation,
C. Yin, “Encoding and decoding dna sequences by in- teger chaos game representation,” Journal of Computa- tional Biology, vol. 26, no. 2, pp. 143–151, 2019. PMID: 30517021
2019
-
[23]
Nucleotide, dinucleotide and trinucleotide frequencies explain patterns observed in chaos game rep- resentations of DNA sequences,
N. Goldman, “Nucleotide, dinucleotide and trinucleotide frequencies explain patterns observed in chaos game rep- resentations of DNA sequences,” Nucleic Acids Research, vol. 21, pp. 2487–2491, 05 1993
1993
-
[25]
Schroeder, Fractals, Chaos, Power Laws: Minutes from an Infinite Paradise
M. Schroeder, Fractals, Chaos, Power Laws: Minutes from an Infinite Paradise . 2012
2012
-
[1496]
would like to thank the support from DGAPA-UNAM under the Project UNAM-PAPIIT IA103325
Also, M.A.Q.-J. would like to thank the support from DGAPA-UNAM under the Project UNAM-PAPIIT IA103325. A.M.E.-R. thanks the support from UAM re- search grant 2024-CPIR-0
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.