Pith. sign in

REVIEW 1 cited by

SNP2Vec: Scalable Self-Supervised Pre-Training for Genome-Wide Association Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.06699 v1 pith:VACBQK6I submitted 2022-04-14 cs.LG cs.AI

classification cs.LGcs.AI
keywords understandingpre-trainingsnp2vecapproachmethodsself-supervisedassociationgenome-wide
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised pre-training methods have brought remarkable breakthroughs in the understanding of text, image, and speech. Recent developments in genomics has also adopted these pre-training methods for genome understanding. However, they focus only on understanding haploid sequences, which hinders their applicability towards understanding genetic variations, also known as single nucleotide polymorphisms (SNPs), which is crucial for genome-wide association study. In this paper, we introduce SNP2Vec, a scalable self-supervised pre-training approach for understanding SNP. We apply SNP2Vec to perform long-sequence genomics modeling, and we evaluate the effectiveness of our approach on predicting Alzheimer's disease risk in a Chinese cohort. Our approach significantly outperforms existing polygenic risk score methods and all other baselines, including the model that is trained entirely with haploid sequences. We release our code and dataset on https://github.com/HLTCHKUST/snp2vec.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BMFM-DNA: A SNP-aware DNA foundation model to capture variant effects

    q-bio.GN 2025-06 conditional novelty 5.0 of 10

    Encoding human genetic variants as special characters during DNA foundation model pre-training is claimed to improve downstream task performance, with small margins and a confounded comparison.

Pith tools