REVIEW 3 major objections 5 minor 5 references
How chromatin interactions shed light on interpreting non-coding genomic variants: opportunities and future direc-tions
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that muscle-disease GWAS variants are enriched in Hi-C-defined TAD domains and chromatin loops, making 3D genome maps a useful layer for interpreting non-coding variants.
desk verdict The review part is a competent but unoriginal tour of 3D genome-variant links; the paper's own enrichment analysis collapses on coordinate mismatch and a missing null. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The core objects are Topologically Associating Domains (TADs), chromatin loops, and A/B compartments called from Hi-C contact maps at 5 kb resolution, plus the overlap analysis that counts GWAS SNPs and CNVs falling inside these features versus matched random regions. TADs are the genome's insulated regulatory neighborhoods; loops are point-to-point enhancer-promoter contacts; compartments are active and inactive nuclear segregation. The argument's workhorse is the enrichment comparison: the observed overlaps (4,098 SNPs in 6,232 TADs; 1,523 SNPs in 268 loops) are set against random-region overlaps (710 and 391, respectively) to show that disease variants concentrate in structured chromatin.
What would settle it
Re-run the overlap analysis after lifting the Hi-C-derived TAD, loop, and compartment coordinates to hg38, or moving all variant coordinates back to hg19. If the enrichment ratios collapse, with the 4,098 vs. 710 TAD hits or the 1,523 vs. 391 loop hits dropping to near-random levels, the central quantitative claim fails. The check is a coordinate conversion and a recount.
Extended reading notes
Core claim
The central claim is that chromatin interaction data can illuminate the regulatory impact of non-coding variants, and that disease-associated variants are not randomly distributed across 3D genome features. Using an in-house muscle Hi-C library, the authors find significant enrichment of muscle-disease GWAS SNPs in TADs and in chromatin loops compared with random genomic regions. Neuromuscular CNVs frequently overlap TADs and other regulatory domains, and the narrative review links such overlaps to known pathogenic mechanisms: boundary disruption can merge or create TADs, loop-altering variants can change enhancer-promoter contacts, and compartment shifts can reposition genes into repressive environments. The intended consequence is that variant interpretation pipelines should treat Hi-C-derived domains and loops as functional annotation layers.
Load-bearing premise
The load-bearing premise is that every set of genomic coordinates being compared sits on the same reference-genome build; the paper aligns Hi-C data to hg19 but lifts variants to hg38, and if the Hi-C features were not also converted, all reported enrichment counts would be suspect.
Editorial extensions
If this is right
- Disease-associated SNPs are significantly overrepresented inside TAD domains and chromatin loops from muscle Hi-C, so Hi-C annotations can prioritize non-coding variants for functional follow-up.
- CNVs that delete or duplicate TAD boundaries or loop anchors can rewire enhancer-promoter contacts and cause misregulation, supporting the use of 3D genome context in variant interpretation.
- Non-coding variants' regulatory impact can be studied by integrating Hi-C with epigenetic marks; variants within loops near enhancers are stronger candidates for affecting gene expression.
- Existing tools that prioritize CNVs for disease associations should incorporate chromatin interactions, especially for non-coding CNV regions.
- Detailed overlap maps of TAD boundaries, chromatin loops, and disease-associated SNPs could reveal evolutionary conservation and tissue-specific regulatory activity.
Reading between the lines
- Inference: If the enrichment survives a genome-build correction, Hi-C maps could be used as a quantitative prior, with disease SNPs inside loops and TADs ranked higher for CRISPR or reporter validation.
- Inference: The loop enrichment (1,523 SNPs in 268 loops vs. 391 in random regions) suggests that loop anchors are variant-dense; testing whether these SNPs coincide with CTCF or cohesin binding motifs would connect the statistical overlap to a mechanistic model.
- Inference: The reported 50 kb boundary-proximal effect implies that distance-to-TAD-boundary could serve as a continuous predictor of variant impact, so an independent test would check whether enrichment rises monotonically toward boundaries.
- Inference: A natural extension of the paper's logic is to ask whether the same enrichment holds across many cell types and diseases, which would determine whether a single Hi-C map can prioritize variants broadly or whether tissue-matched maps are required.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript, framed as a review with in-house analyses, argues that integrating chromatin interaction data (Hi-C) with non-coding CNV and GWAS variant data can help interpret the regulatory impact of non-coding variants. The authors used in-house Hi-C data from a muscle cell line, called TADs, loops, and compartments, and then quantified overlaps of these features with neuromuscular-disease CNVs and muscle-disease GWAS SNPs. The central quantitative claims are that 4,098 muscle-disease GWAS SNPs overlap 6,232 TADs (versus 710 in random regions) and that 1,523 SNPs overlap 268 loops (versus 391 in random regions), which the authors present as significant enrichment. The paper also reviews literature on TAD disruption, loops, compartments, enhancer-promoter interactions, and super-enhancers in disease.
Significance. The topic is timely and clinically relevant: methods to prioritize non-coding variants are needed, and Hi-C-based features could in principle provide such prioritization. The literature review is broad and covers many relevant studies. The in-house analysis, however, is the part that purports to provide new quantitative evidence, and that evidence is currently not statistically substantiated. If the overlap counts were computed correctly on a consistent genome build and compared against a well-defined null, they could support the paper's thesis. As written, the enrichment claim is uninterpretable because the genome builds of the variant and Hi-C feature sets may be mismatched and no formal statistical test or null-model description is supplied. The paper's value therefore rests mostly on its review component, which is competent but not novel.
major comments (3)
- [Methods, Preparation of Hi-C Libraries; Methods, Genomic Variants] The Hi-C data are explicitly aligned to hg19, and TADs, loops, and compartments are derived from that hg19 map. In contrast, the CNV table (Table 1) reports hg38 coordinates after liftover from hg19, and the genome build of the GWAS SNP files is never stated. The overlap analyses in the Results (e.g., the 4,098/710 TAD enrichment and 1,523/391 loop enrichment) therefore appear to intersect hg19 features with hg38 variants, or with variants of unknown build. If the builds are inconsistent, every overlap count in the paper is meaningless. The authors must state the build of every dataset, ensure all variants and features are on the same build (e.g., lift over the Hi-C features or re-align to hg38), and recompute all overlap statistics after this correction.
- [Results, Disrupting topologically associating domain boundaries; Results, Disrupting chromatin loops] No statistical test is reported for the enrichment counts. The text says the SNP overlap is 'significantly higher' and 'significant enrichment,' but no p-value, confidence interval, or effect-size measure is given. The random-region control is described only as 'an equivalent number of randomly generated regions' without specifying the number of random regions, their size distribution, whether they exclude unmappable or blacklisted regions, whether they are matched for SNP density, GC content, or gene content, or how many randomizations were performed. Without a defined null model and a formal test, the counts 4,098 vs 710 and 1,523 vs 391 cannot be taken as evidence of enrichment. This is a load-bearing issue because the enrichment claim is the main in-house quantitative contribution.
- [Results, Disrupting topologically associating domain boundaries] The statement that SNPs within 50 kb of TAD boundaries 'were significantly more likely to affect gene expression through long-range regulatory interactions' is not supported by any analysis shown in the manuscript. No comparison is provided between boundary-proximal and boundary-distal SNPs, no effect measure is defined, and no test result is reported. Either the analysis should be presented with full statistical detail, or the claim should be removed or clearly attributed to the cited literature rather than to the authors' own results.
minor comments (5)
- [Throughout] The manuscript contains numerous typos and grammatical errors that impede readability, for example 'repsresented,' 'protray,' 'abnoramlity,' 'polimorphisms,' 'disoreder,' 'to to,' and inconsistent punctuation such as double periods. A thorough language edit is needed.
- [Methods, Preparation of Hi-C Libraries] The interaction callers MaxHiC and MHiC are developed by the authors' own group. This is not inherently problematic, but because the enrichment analysis depends entirely on these callers, the authors should state this provenance explicitly and ideally provide a comparison with or validation against an independent caller (e.g., FitHiC2 or HICCUPS) to rule out caller-specific artifacts.
- [Methods, Genomic Variants] Two of the 79 CNVs were excluded because they 'mapped to multiple regions in the new build,' but the exclusion criterion is not described in detail. It is unclear whether multi-mapping was defined at the level of the entire CNV or individual liftover blocks, and whether this exclusion could bias the overlap results. Please clarify.
- [Results, Alterations in chromatin interaction maps] Figure 1 is described as showing chromosome 1 and chromosome 14 contact maps, but the text does not state which cell line or condition the maps come from, nor how the '40k resolution' heatmaps were generated or normalized. Adding this information would improve reproducibility.
- [Results, Disrupting chromatin loops] The loop analysis reports 2,582 loops in the 'muscular tissue library RT,' but the methods do not define what RT, 27M, and 26F refer to (replicates? individuals? conditions?). Figure 6 lists 'RT, 27M, 26F' without explanation. Please define these labels.
Circularity Check
No significant circularity: the enrichment claims are empirical overlap counts, and the self-citations to in-house Hi-C callers are not load-bearing in a circular sense.
full rationale
The paper's central quantitative claims are empirical overlaps between externally sourced GWAS/CNV coordinates and Hi-C-derived TADs, loops, and compartments. Nothing in the derivation defines the Hi-C features in terms of the variant data or vice versa: the TAD and loop calls come from MaxHiC/MHiC and hicExplorer applied to in-house muscle Hi-C libraries, while the GWAS SNPs come from the GWAS Catalog and GWASdb and the CNVs from a published exome study. The reported counts (e.g., 4,098 SNPs in 6,232 TADs vs. 710 in random regions; 1,523 SNPs in 268 loops vs. 391 in random regions) are overlap statistics, not quantities fitted from those same overlaps, so there is no reduction of a 'prediction' to its inputs. The self-citations to MaxHiC [27] and MHiC [28] are tool citations rather than imported theorems: those tools are separately published, code-based methods whose stated assumptions concern Hi-C background correction and do not include the muscle-disease GWAS enrichment result. Earlier self-citations in the Introduction ([10-16]) are context about the authors' prior work and are not load-bearing for the derivation. The genuine weaknesses of the paper are validity/reproducibility problems, not circularity: no statistical test or P-value accompanies the 'significantly higher' enrichment statements, the random-region generation is not described, and the coordinate-build consistency is questionable because the Hi-C pipeline is explicitly hg19 while the CNV table is hg38 with no stated liftover for Hi-C features. These issues undermine the strength of the conclusions, but they do not make the derivation equivalent to its inputs by construction.
Assumptions & free parameters
free parameters (4)
- Hi-C interaction significance thresholds =
P<0.01, read count >=10, distance 5kb-10Mb
- Promoter overlap threshold =
>=10% overlap
- TAD boundary FDR threshold =
FDR <0.05
- GWAS phenotype filter for muscle disease =
Not stated
assumptions (4)
- domain assumption Genome assembly coordinate consistency
- domain assumption Random regions form an appropriate null model
- domain assumption In-house muscle Hi-C library is representative of muscle regulatory architecture
- standard math LiftOver accurately maps CNVs between hg19 and hg38
Cite this review
Pith. "Pith review of How chromatin interactions shed light on interpreting non-coding genomic variants: opportunities and future direc-tions." pith.science (2026). https://pith.science/paper/MKWXY2W7
@misc{pith2026241117956,
author = {Pith},
title = {Pith review of: How chromatin interactions shed light on interpreting non-coding genomic variants: opportunities and future direc-tions},
year = {2026},
howpublished = {\url{https://pith.science/paper/MKWXY2W7}},
note = {Machine review of arXiv:2411.17956}
}
read the original abstract
Genomic variants, including copy number variants (CNVs) and genome-wide associa-tion study (GWAS) single nucleotide polymorphisms (SNPs), represent structural alterations that influence genomic diversity and disease susceptibility. While coding region variants have been extensively studied, non-coding and regulatory variants present significant challenges due to their potential impacts on gene regulation, which are often obscured by the complexity of the ge-nome. Chromatin interactions, which organize the genome spatially and regulate gene expression through enhancer-promoter contacts, predominantly occur in non-coding regions. Notably, more than 90% of enhancers, crucial for gene regulation, reside in these non-coding regions, underscor-ing their importance in interpreting the regulatory effects of CNVs and GWAS-associated SNPs. In this study, we integrate chromatin interaction data with CNV and GWAS data to uncover the functional implications of non-coding variants. By leveraging this integrated approach, we pro-vide new insights into how structural variants and disease-associated SNPs disrupt regulatory networks, advancing our understanding of genetic complexity. These findings offer potential av-enues for personalized medicine by elucidating disease mechanisms and guiding therapeutic strategies tailored to individual genomic profiles. This research underscores the critical role of chromatin interactions in revealing the regulatory consequences of non-coding variants, bridging the gap between genetic variation and phenotypic outcomes.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[22]
Krijger, P.H. and W. de Laat, Regulation of disease-associated gene expression in the 3D genome. Nat Rev Mol Cell Biol, 2016. 17(12): p. 771-782. 23. Lupiáñez, D.G., et al., Disruptions of topological chromatin domains cause pathogenic rewiring of gene-enhancer interactions. Cell, 2015. 161(5): p. 1012-1025. 24. Spielmann, M. and S. Mundlos, Structural va...
work page 2016
-
[47]
Dekker, J., M.A. Marti-Renom, and L.A. Mirny, Exploring the three-dimensional organization of genomes: interpreting chromatin interaction data. Nat Rev Genet, 2013. 14(6): p. 390-403. 48. Nakato, R., et al., Context-dependent perturbations in chromatin folding and the transcriptome by cohesin and related factors. Nat Commun, 2023. 14(1): p. 5647. 49. Li, ...
work page 2013
-
[70]
Tan, W.L.W., et al., Epigenomes of Human Hearts Reveal New Genetic Variants Relevant for Cardiac Disease and Phenotype. Circulation Research, 2020. 127(6): p. 761-777. 71. Yuan, X., I.C. Scott, and M.D. Wilson, Heart Enhancers: Development and Disease Control at a Distance. Front Genet, 2021. 12: p. 642975. 72. Tsang, Felice H., et al., The characteristic...
work page 2020
-
[93]
Peric-Hupkes, D., et al., Molecular maps of the reorganization of genome-nuclear lamina interactions during differentiation. Mol Cell, 2010. 38(4): p. 603-13. 94. Tai, P.W., et al., The dynamic architectural and epigenetic nuclear landscape: developing the genomic almanac of biology and disease. J Cell Physiol, 2014. 229(6): p. 711-27. 95. Norton, H.K. an...
work page 2010
-
[118]
Vahedi, G., et al., Super-enhancers delineate disease-associated regulatory nodes in T cells. Nature, 2015. 520(7548): p. 558-562. 119. Li, X.P., et al., The Emerging Role of Super-enhancers as Therapeutic Targets in The Digestive System Tumors. Int J Biol Sci, 2023. 19(4): p. 1036-1048. 120. Rahaie, Z., Rabiee, H. R., & Alinejad-Rokny, H., DeepGenePrior:...
work page 2015
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.