Pretraining a small DNA model on unlabeled task-related sequences improves gene finding and other BEND tasks compared to training from scratch, with less data than genome-scale pretraining.
Never train from scratch: Fair comparison of long-sequence models requires data-driven priors
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
q-bio.GN 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Improving Genomic Models via Task-Specific Self-Pretraining
Pretraining a small DNA model on unlabeled task-related sequences improves gene finding and other BEND tasks compared to training from scratch, with less data than genome-scale pretraining.