{"id":"e6dda018-5fdf-4dc0-a8a3-6fec92ef5272","arxiv_id":"2411.17956","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A mostly review preprint arguing Hi-C data can help interpret non-coding variants, with a small enrichment analysis that is incompletely described and likely affected by a genome build mismatch.","lead":"This paper pairs a literature review of how 3D chromatin architecture helps interpret non-coding DNA changes with a small, loosely documented overlap analysis between muscle disease variants and Hi-C derived regions. It offers a generalist-friendly entry point into non-coding variant interpretation, but the new analysis is not sound as reported.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central enrichment claim lacks a valid comparison: Hi-C features are hg19, CNVs are hg38 and GWAS build is unstated, and the random-region null plus statistical test are undefined. Matched-coordinate and matched-null recomputation is required before the counts can support significance.","rationale":"The review sections are a reasonable synthesis of the existing literature on Hi-C, TADs, loops, compartments, and variant interpretation, and I am not disputing that portion. The problem is that the paper's claimed new contribution—significant enrichment of muscle-disease GWAS SNPs in TADs and loops—rests entirely on the overlap analysis in the sections 'Disrupting topologically associating domain boundaries' and 'Disrupting chromatin loops'. The methods for that analysis are incomplete in exactly the places that determine the numbers. The build mismatch is the sharpest symptom: Hi-C/TAD/loop features are stated to be hg19, while the CNV table is hg38 and the GWAS SNP build is never stated; if the SNP files come from the current GWAS Catalog, their positions would be GRCh38, making the overlap with hg19 features invalid. Equally damaging, the random-region comparison is not described well enough to know whether it controls for any confounder; 'equivalent number' says nothing about size matching, mappability, or SNP density, and no test statistic is given. These are internal consistency and reproducibility problems, not differences of biological opinion. A concrete check—re-running the pipeline from the linked repository after unifying builds and specifying the null—would settle whether the enrichment is real. As the manuscript stands, the quantitative conclusion is not supported, so rejection is warranted; if a reanalysis fixes these points, the review portion could stand alone or the analysis could be re-added with proper coordinate unification and statistical controls.","tokens_in":16685,"tokens_out":6146,"duration_ms":57478,"concrete_test":"Access the repository at https://github.com/jade0530/hic_review; inspect the exact GWAS input files and their build, then unify coordinates (lift GWAS/CNV variants to hg19 or lift Hi-C TAD/loop calls to hg38) and recompute all overlaps. Next regenerate random regions with the same count, size distribution, chromosome coverage, and mappability/GC constraints as the TAD/loop sets, and compute enrichment with a stated test (e.g., logistic regression controlling for SNP density). If the counts shift or the effect disappears, the paper's central claim is unsupported.","verdict_should_be":"REJECT","load_bearing_attack":"To support the claim that muscle-disease GWAS SNPs are enriched in TADs and loops, the overlap analysis needs (a) variant and feature coordinates on the same genome build, and (b) a valid null. Condition (a) is not met as written: the Hi-C pipeline is hg19 (Methods, 'Preparation of Hi-C Libraries'), TAD/loop calls derive from that hg19 map, yet the CNV Table 1 is explicitly hg38 and no build or liftover is stated for the GWAS SNP files before intersecting them with hg19 features. Condition (b) is also missing: the random-region controls are only 'equivalent number' of random regions; their count, length distribution, exclusion of unmappable/blacklisted regions, and matching for SNP density/GC are not described, and no statistical test, p-value, or confidence interval is reported. The reported counts (4,098 vs 710 for TADs; 1,523 vs 391 for loops) are therefore uninterpretable as evidence of enrichment. This is an internal reproducibility failure, not a disagreement with field consensus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript, framed as a review with in-house analyses, argues that integrating chromatin interaction data (Hi-C) with non-coding CNV and GWAS variant data can help interpret the regulatory impact of non-coding variants. The authors used in-house Hi-C data from a muscle cell line, called TADs, loops, and compartments, and then quantified overlaps of these features with neuromuscular-disease CNVs and muscle-disease GWAS SNPs. The central quantitative claims are that 4,098 muscle-disease GWAS SNPs overlap 6,232 TADs (versus 710 in random regions) and that 1,523 SNPs overlap 268 loops (versus 391 in random regions), which the authors present as significant enrichment. The paper also reviews literature on TAD disruption, loops, compartments, enhancer-promoter interactions, and super-enhancers in disease.","tokens_in":16820,"tokens_out":2683,"duration_ms":25504,"significance":"The topic is timely and clinically relevant: methods to prioritize non-coding variants are needed, and Hi-C-based features could in principle provide such prioritization. The literature review is broad and covers many relevant studies. The in-house analysis, however, is the part that purports to provide new quantitative evidence, and that evidence is currently not statistically substantiated. If the overlap counts were computed correctly on a consistent genome build and compared against a well-defined null, they could support the paper's thesis. As written, the enrichment claim is uninterpretable because the genome builds of the variant and Hi-C feature sets may be mismatched and no formal statistical test or null-model description is supplied. The paper's value therefore rests mostly on its review component, which is competent but not novel.","major_comments":[{"comment":"The Hi-C data are explicitly aligned to hg19, and TADs, loops, and compartments are derived from that hg19 map. In contrast, the CNV table (Table 1) reports hg38 coordinates after liftover from hg19, and the genome build of the GWAS SNP files is never stated. The overlap analyses in the Results (e.g., the 4,098/710 TAD enrichment and 1,523/391 loop enrichment) therefore appear to intersect hg19 features with hg38 variants, or with variants of unknown build. If the builds are inconsistent, every overlap count in the paper is meaningless. The authors must state the build of every dataset, ensure all variants and features are on the same build (e.g., lift over the Hi-C features or re-align to hg38), and recompute all overlap statistics after this correction.","section":"Methods, Preparation of Hi-C Libraries; Methods, Genomic Variants"},{"comment":"No statistical test is reported for the enrichment counts. The text says the SNP overlap is 'significantly higher' and 'significant enrichment,' but no p-value, confidence interval, or effect-size measure is given. The random-region control is described only as 'an equivalent number of randomly generated regions' without specifying the number of random regions, their size distribution, whether they exclude unmappable or blacklisted regions, whether they are matched for SNP density, GC content, or gene content, or how many randomizations were performed. Without a defined null model and a formal test, the counts 4,098 vs 710 and 1,523 vs 391 cannot be taken as evidence of enrichment. This is a load-bearing issue because the enrichment claim is the main in-house quantitative contribution.","section":"Results, Disrupting topologically associating domain boundaries; Results, Disrupting chromatin loops"},{"comment":"The statement that SNPs within 50 kb of TAD boundaries 'were significantly more likely to affect gene expression through long-range regulatory interactions' is not supported by any analysis shown in the manuscript. No comparison is provided between boundary-proximal and boundary-distal SNPs, no effect measure is defined, and no test result is reported. Either the analysis should be presented with full statistical detail, or the claim should be removed or clearly attributed to the cited literature rather than to the authors' own results.","section":"Results, Disrupting topologically associating domain boundaries"}],"minor_comments":[{"comment":"The manuscript contains numerous typos and grammatical errors that impede readability, for example 'repsresented,' 'protray,' 'abnoramlity,' 'polimorphisms,' 'disoreder,' 'to to,' and inconsistent punctuation such as double periods. A thorough language edit is needed.","section":"Throughout"},{"comment":"The interaction callers MaxHiC and MHiC are developed by the authors' own group. This is not inherently problematic, but because the enrichment analysis depends entirely on these callers, the authors should state this provenance explicitly and ideally provide a comparison with or validation against an independent caller (e.g., FitHiC2 or HICCUPS) to rule out caller-specific artifacts.","section":"Methods, Preparation of Hi-C Libraries"},{"comment":"Two of the 79 CNVs were excluded because they 'mapped to multiple regions in the new build,' but the exclusion criterion is not described in detail. It is unclear whether multi-mapping was defined at the level of the entire CNV or individual liftover blocks, and whether this exclusion could bias the overlap results. Please clarify.","section":"Methods, Genomic Variants"},{"comment":"Figure 1 is described as showing chromosome 1 and chromosome 14 contact maps, but the text does not state which cell line or condition the maps come from, nor how the '40k resolution' heatmaps were generated or normalized. Adding this information would improve reproducibility.","section":"Results, Alterations in chromatin interaction maps"},{"comment":"The loop analysis reports 2,582 loops in the 'muscular tissue library RT,' but the methods do not define what RT, 27M, and 26F refer to (replicates? individuals? conditions?). Figure 6 lists 'RT, 27M, 26F' without explanation. Please define these labels.","section":"Results, Disrupting chromatin loops"}],"recommendation":"major_revision","confidential_remarks":"The paper is a review with an in-house analysis. The central analytical claims currently rest on a likely genome-build mismatch and an undefined statistical null, which are fixable within the manuscript's scope by recomputing the overlaps on a consistent build and adding proper controls. However, if the authors cannot or do not provide the corrected analysis, the quantitative contribution should be removed and the paper treated as a pure review. I also note a heavy reliance on the authors' own tools and references; this is worth checking but does not by itself invalidate the review component."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline is simple: the review material is fine for someone new to the field, but the paper's one original analysis—the SNP enrichment in TADs and loops—does not hold up. The Hi-C features are called on hg19, while the CNVs are explicitly lifted to hg38 and the GWAS build is never stated; without a liftover step for the features, every overlap count in the Results is meaningless. The random-region comparison is equally underspecified: no count of regions, no length distribution, no masking, no matching, and no test statistic or confidence interval. The reader's stress-test is right on both counts.\n\nWhat the paper does well: the narrative review of mechanisms—TAD disruption, loop perturbation, compartment switching, enhancer-promoter contacts—is broadly cited and mostly accurate. The figures, especially the multi-track view, are nice illustrations of how Hi-C can be layered with variant data. If the analysis were redone on a matched genome build with a proper null, the review would be a usable entry point for non-specialists.\n\nSoft spots beyond the load-bearing analysis: there are several typos and stylistic slips, and the heavy reliance on the authors' own Hi-C callers (MaxHiC, MHiC) is not itself a flaw, but it would help to have an independent caller or at least some validation that the loop/TAD calls are reproducible. The selection criteria for 'muscle disease GWAS SNPs' are never given—which phenotype entries were included, and how many SNPs survived QC? That is needed for any interpretation of the counts.\n\nBottom line: I would not cite this as a source for the enrichment claim, and I would not send it to peer review in its current form. The review content is not novel enough to justify referee time on its own, and the quantitative claim is broken. If the authors fix the coordinate issue, add a proper null and a statistical test, and specify the SNP filter, the review portion could be a useful educational piece—but that is a major revision, not a polish.","headline":"The review part is a competent but unoriginal tour of 3D genome-variant links; the paper's own enrichment analysis collapses on coordinate mismatch and a missing null.","tokens_in":17456,"tokens_out":2620,"would_cite":false,"duration_ms":24249,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that muscle-disease GWAS variants are enriched in Hi-C-defined TAD domains and chromatin loops, making 3D genome maps a useful layer for interpreting non-coding variants.","keywords":["non-coding variants","chromatin interactions","Hi-C","GWAS SNPs","copy number variants","topologically associating domains","chromatin loops","gene regulation"],"falsifier":"Re-run the overlap analysis after lifting the Hi-C-derived TAD, loop, and compartment coordinates to hg38, or moving all variant coordinates back to hg19. If the enrichment ratios collapse, with the 4,098 vs. 710 TAD hits or the 1,523 vs. 391 loop hits dropping to near-random levels, the central quantitative claim fails. The check is a coordinate conversion and a recount.","tokens_in":16379,"feed_emoji":"🧬","tokens_out":4929,"duration_ms":41534,"temperature":0.7,"pith_summary":"This review argues that non-coding variants are under-interpreted because standard annotation focuses on coding effects, and that chromatin interaction maps, especially Hi-C, supply the missing regulatory context. To support this, the authors analyze their own muscle Hi-C data and report that GWAS SNPs linked to muscle disease are significantly enriched in TAD domains (4,098 SNPs in 6,232 TADs vs. 710 in random regions) and chromatin loops (1,523 SNPs in 268 loops vs. 391 in random regions). They also survey evidence that CNVs disrupt TAD boundaries, loops, compartments, promoter-enhancer contacts, and super-enhancers. If the enrichment is genuine, Hi-C data becomes a practical filter for prioritizing non-coding variants in disease studies.","feed_headline":"Muscle-disease SNPs cluster inside 3D genome loops","feed_subtitle":"Hi-C maps show 4,098 disease SNPs inside TADs versus 710 in random regions, pointing to non-coding variants worth prioritizing.","key_machinery":"The core objects are Topologically Associating Domains (TADs), chromatin loops, and A/B compartments called from Hi-C contact maps at 5 kb resolution, plus the overlap analysis that counts GWAS SNPs and CNVs falling inside these features versus matched random regions. TADs are the genome's insulated regulatory neighborhoods; loops are point-to-point enhancer-promoter contacts; compartments are active and inactive nuclear segregation. The argument's workhorse is the enrichment comparison: the observed overlaps (4,098 SNPs in 6,232 TADs; 1,523 SNPs in 268 loops) are set against random-region overlaps (710 and 391, respectively) to show that disease variants concentrate in structured chromatin.","core_discovery":"The central claim is that chromatin interaction data can illuminate the regulatory impact of non-coding variants, and that disease-associated variants are not randomly distributed across 3D genome features. Using an in-house muscle Hi-C library, the authors find significant enrichment of muscle-disease GWAS SNPs in TADs and in chromatin loops compared with random genomic regions. Neuromuscular CNVs frequently overlap TADs and other regulatory domains, and the narrative review links such overlaps to known pathogenic mechanisms: boundary disruption can merge or create TADs, loop-altering variants can change enhancer-promoter contacts, and compartment shifts can reposition genes into repressive environments. The intended consequence is that variant interpretation pipelines should treat Hi-C-derived domains and loops as functional annotation layers.","pith_inferences":["Inference: If the enrichment survives a genome-build correction, Hi-C maps could be used as a quantitative prior, with disease SNPs inside loops and TADs ranked higher for CRISPR or reporter validation.","Inference: The loop enrichment (1,523 SNPs in 268 loops vs. 391 in random regions) suggests that loop anchors are variant-dense; testing whether these SNPs coincide with CTCF or cohesin binding motifs would connect the statistical overlap to a mechanistic model.","Inference: The reported 50 kb boundary-proximal effect implies that distance-to-TAD-boundary could serve as a continuous predictor of variant impact, so an independent test would check whether enrichment rises monotonically toward boundaries.","Inference: A natural extension of the paper's logic is to ask whether the same enrichment holds across many cell types and diseases, which would determine whether a single Hi-C map can prioritize variants broadly or whether tissue-matched maps are required."],"forward_implications":["Disease-associated SNPs are significantly overrepresented inside TAD domains and chromatin loops from muscle Hi-C, so Hi-C annotations can prioritize non-coding variants for functional follow-up.","CNVs that delete or duplicate TAD boundaries or loop anchors can rewire enhancer-promoter contacts and cause misregulation, supporting the use of 3D genome context in variant interpretation.","Non-coding variants' regulatory impact can be studied by integrating Hi-C with epigenetic marks; variants within loops near enhancers are stronger candidates for affecting gene expression.","Existing tools that prioritize CNVs for disease associations should incorporate chromatin interactions, especially for non-coding CNV regions.","Detailed overlap maps of TAD boundaries, chromatin loops, and disease-associated SNPs could reveal evolutionary conservation and tissue-specific regulatory activity."],"supporting_citations":[{"why":"Supplies the neuromuscular CNV list used in the overlap and TAD visualization analysis.","marker":"[35]"},{"why":"Source of GWAS SNP data for the enrichment analysis.","marker":"[32, 33]"},{"why":"HiC-Pro processed and aligned the in-house Hi-C data at 5 kb resolution.","marker":"[26]"},{"why":"MaxHiC identifies statistically significant Hi-C interactions used in the analysis.","marker":"[27]"},{"why":"MHiC provides additional identification and visualization of significant interactions.","marker":"[28]"},{"why":"hicExplorer called TADs and loops from the Hi-C matrices.","marker":"[38-40]"},{"why":"HiTC detected A/B compartments used to characterize active and inactive regions.","marker":"[41]"},{"why":"Landmark evidence that TAD boundary disruptions rewire enhancer-promoter interactions and cause pathogenic phenotypes.","marker":"[23]"},{"why":"Kilobase-resolution 3D map establishing CTCF and cohesin loops, framing the loop analysis.","marker":"[19]"},{"why":"Supports the claim that TAD boundaries are required for normal genome function and that boundary-proximal SNPs affect gene expression.","marker":"[64]"}],"fun_headline_variants":["Disease variants cluster in 3D genome loops and TADs","Chromatin folds guide interpretation of non-coding variants","3D genome maps spotlight regulatory impact of disease SNPs","TAD boundaries and loops help prioritize non-coding mutations","Muscle-disease SNPs enriched in 3D genome TADs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every set of genomic coordinates being compared sits on the same reference-genome build; the paper aligns Hi-C data to hg19 but lifts variants to hg38, and if the Hi-C features were not also converted, all reported enrichment counts would be suspect.","fun_headline_variants_meta":{"raw":{"variants":["Disease variants cluster in 3D genome loops and TADs","Chromatin folds guide interpretation of non-coding variants","3D genome maps spotlight regulatory impact of disease SNPs","TAD boundaries and loops help prioritize non-coding mutations","Muscle-disease SNPs enriched in 3D genome TADs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000283,"raw_usage":{"total_tokens":1665,"prompt_tokens":933,"completion_tokens":732,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":663}},"tokens_in":549,"tokens_out":732,"duration_ms":6877,"temperature":1.0,"reasoning_tokens":663,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:39:29.030705+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the overlap analysis after lifting the Hi-C-derived TAD, loop, and compartment coordinates to hg38, or moving all variant coordinates back to hg19. If the enrichment ratios collapse, with the 4,098 vs. 710 TAD hits or the 1,523 vs. 391 loop hits dropping to near-random levels, the central quantitative claim fails. The check is a coordinate conversion and a recount.","supporting_citations":[],"review_version":1}