REVIEW 4 cited by
Faster and More Accurate Sequence Alignment with SNAP
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
We present the Scalable Nucleotide Alignment Program (SNAP), a new short and long read aligner that is both more accurate (i.e., aligns more reads with fewer errors) and 10-100x faster than state-of-the-art tools such as BWA. Unlike recent aligners based on the Burrows-Wheeler transform, SNAP uses a simple hash index of short seed sequences from the genome, similar to BLAST's. However, SNAP greatly reduces the number and cost of local alignment checks performed through several measures: it uses longer seeds to reduce the false positive locations considered, leverages larger memory capacities to speed index lookup, and excludes most candidate locations without fully computing their edit distance to the read. The result is an algorithm that scales well for reads from one hundred to thousands of bases long and provides a rich error model that can match classes of mutations (e.g., longer indels) that today's fast aligners ignore. We calculate that SNAP can align a dataset with 30x coverage of a human genome in less than an hour for a cost of $2 on Amazon EC2, with higher accuracy than BWA. Finally, we describe ongoing work to further improve SNAP.
Forward citations
Cited by 4 Pith papers
-
The anti-lexicographic SUS-anchor: a near-optimal k=1 sampling scheme
The anti-lexicographic SUS-anchor achieves sampling densities less than 1% above the lower bound for alphabet size 4 and k=1, substantially outperforming bidirectional anchors.
-
Inductive-bias-driven Reinforcement Learning For Efficient Schedules in Heterogeneous Clusters
Symphony uses a domain-driven Bayesian network as an inductive bias in an RL scheduler, dramatically cutting training data needs while beating black-box methods.
-
Extending TensorFlow's Semantics with Pipelined Execution
PTF adds stages, gates, and per-feed metadata to TensorFlow to support concurrent, isolated, flow-controlled processing of multiple batches, demonstrated on a genomic align/sort pipeline.
-
Dependencies and Dataflow in Seed-Filter-Extend Pipelines
The paper analyzes dependencies in genome alignment pipelines and implements synthesized optimizations from four prior tools into LASTZ to reduce serial bottlenecks.
Discussion (0). Continue with ORCID to comment.