REVIEW 3 major objections 7 minor 94 references
Contributions of the Petabyte Scale Sequence Search Codeathon toward efforts to scale sequence-based searches on SRA
T0 review · 3 major / 7 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Four codeathon teams produced reusable benchmarks and proof-of-concept pipelines for petabyte-scale sequence search, the paper reports.
desk verdict An honestly scoped codeathon report with real, reusable artifacts; the BLAST-as-gold-standard circularity is genuine but does not sink its central claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a set of benchmark blueprints: a gold-standard definition of sequence containment; the use of BLASTn as ground truth; precision/recall evaluation against that ground truth; workflow layers built with Snakemake and Nextflow; the BAGEL harness for hard-annotation benchmarks; and the two-stage Pebblescout-to-ElasticBLAST search. The containment definition and BLAST ground truth carry the evaluation numbers, the workflow frameworks carry reproducibility, and the two-stage pipeline carries the demonstration that an index can narrow a petabyte-scale corpus to a few hundred plausible candidates before alignment.
What would settle it
Take a metagenome sample, spike in contigs at known identities and lengths, then run the gold-standard construction and count how many spiked containments BLASTn misses and how many unrelated pairs it reports; if BLASTn disagrees with the spike-in truth by more than a small margin, the precision and recall rankings in the benchmarks do not reflect true search accuracy.
Extended reading notes
Core claim
The codeathon's contribution is a decomposition of petabyte-scale sequence search into benchmarkable subproblems, each supported by a working prototype. Team 1's bothie workflow wraps several k-mer sketching tools and uses BLAST-based comparisons as likely ground truth to detect long query sequences inside metagenomes. Team 2 defined containment as an alignment with more than 95% identity covering more than 95% of the shorter contig, generated gold standards by running BLASTn, and built a Snakemake pipeline that scores other tools by precision and recall stratified by identity and contig length. Team 3 produced pilot hard-annotation benchmarks and the BAGEL harness for running them on local, cluster, and cloud environments. Team 4 combined the Pebblescout k-mer index with ElasticBLAST to prefilter candidate SRA runs and then align queries against them, finding that the main bottleneck was downloading the selected reads rather than the search itself. Together these proofs-of-concept, with the datasets and repositories shared alongside them, are the paper's contribution.
Load-bearing premise
The benchmarking results assume that BLASTn's alignments are a correct and complete gold standard for DNA containment, so any match BLAST misses or falsely reports is inherited by every tool's precision and recall numbers.
Editorial extensions
If this is right
- Researchers can reproduce and extend the codeathon benchmarks from public repositories rather than reimplementing evaluation from scratch.
- A biologist with a query genome or biosynthetic gene cluster can use the bothie pipeline to screen metagenomes for near-identical containing samples.
- The Team 2 gold standards provide a baseline against which new containment tools can be scored by precision and recall, stratified by identity and contig length.
- The Team 4 pipeline shows that after a k-mer index reduces the archive to candidate runs, the practical limit becomes download time rather than search time.
- The BAGEL harness lets tool developers add a new alignment or annotation tool by wrapping it in a container, without writing cluster or cloud code.
Reading between the lines
- Inference: because the pipelines and repositories are public, a corrected gold standard could be swapped in without rebuilding the evaluation framework, so the infrastructure's value does not depend on BLAST being perfect.
- Inference: the narrow containment definition (more than 95% identity over more than 95% of the shorter contig) excludes partial or diverged matches; rerunning the benchmarks with a looser definition may change which tools look best for viral-discovery use cases.
- Inference: Team 2's proposed order-preserving minimizer sketch, if completed into a colored de Bruijn graph, could make containment queries nearly constant-time and is a natural next step to test within the same harness.
- Inference: the download bottleneck identified by Team 4 suggests that cloud-native formats and direct streaming from object storage, rather than faster alignment, are the near-term limiting factor for SRA-wide pipelines.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports on a virtual codeathon held September 27 - October 1, 2021, convened by DOE BER, NIH ODSS, and NCBI to develop benchmarking approaches and reusable software for petabyte-scale sequence-based search of the NCBI Sequence Read Archive, with a focus on metagenomic data. Four teams produced early-stage deliverables: Team 1 developed 'bothie,' a pipeline for detecting user-provided long query sequences with k-mer-based tools using BLAST as a 'likely groundtruth'; Team 2 developed a containment-benchmark pipeline with BLASTn as the gold standard and evaluated MMseqs2, Minimap2, dashing, MashMap, and a synteny-based method; Team 3 developed the BAGEL harness and pilot hard-annotation benchmarks, and Team 4 built a combined Pebblescout + ElasticBLAST pipeline. The Discussion claims the codeathon successfully produced benchmarks, reusable software, and a cloud-based foundation for scaling SRA searches. The paper is written as a codeathon report with modest, explicitly preliminary claims, and it emphasizes reproducibility through public repositories and workflow tools.
Significance. The paper's strengths are its documentation of a structured community effort, the public availability of the four team repositories (bothie, psss-team2, psss-team3-hard-annotation, psss-team4), and the candid description of practical bottlenecks such as the fasterq-dump transfer problem. As an organizational and infrastructure contribution, the manuscript has value for future codeathon organizers and for researchers seeking reusable scaffolding for metagenomic search workflows. However, the central claim of producing 'benchmarks' is not backed by reported evaluation results: no precision/recall numbers, error bars, or statistical analyses appear in the text. The most load-bearing assumption, BLAST as gold standard while BLAST itself is a benchmarked tool, makes the Team 2 evaluation circular and limits the value of the claimed benchmark as an independent resource. The paper would be strengthened by softening the 'benchmark' framing or by presenting actual metrics and an accompanying validation strategy.
major comments (3)
- [§5.2.2, Table 2] The benchmarking methodology in Section 5.2.2 treats BLASTn output as the gold standard for containment ('The results of using BLASTn for the search were treated as the gold standard'), while BLAST is itself one of the tools evaluated in Table 2. This makes the precision/recall calculations circular: every true positive, false positive, and false negative for MMseqs2, Minimap2, dashing, MashMap, and synteny is defined relative to BLAST's own sensitivity and false-positive profile. A containment missed by BLAST is scored as a false negative for the other tools, and a spurious BLAST hit is scored as ground truth. This undermines the claim in Section 6 that the codeathon produced 'benchmarks and reusable software' as an evaluation resource, because the benchmark cannot independently rank methods whose errors are defined by the very tool being assessed. I recommend either (i) constructing ground truth from simulated reads/contigs with known containment, (ii) using a consensus of several independent aligners, or (iii) clearly relabeling all reported precision/recall as 'agreement with BLAST' rather than as absolute accuracy.
- [§5.2.2 and §5.2.3] The text states that precision and recall were computed for the stool metagenome dataset ('Containments were identified using all vs all alignments to obtain precision and recall values on BLAST, Dashing, MiniMap2 and MashMap'), and Section 5.2.3 describes TP/FP/FN classification and stratification by percent identity, contig length, and confidence. However, no precision, recall, TP/FP/FN, or stratified values are reported anywhere in the paper. The only quantitative results are aggregate containment counts in Section 5.2.2 (155,372 total containments; >141k BLAST/MMseq2 agreement; 75k synteny overlap). Since the central deliverable is a benchmark, the absence of the actual benchmark scores makes the results non-reproducible as stated and weakens the Section 6 claim of 'rapidly producing benchmarks.' Please add tables or figures with the precision/recall and runtime/memory measurements mentioned in Section 5.2.3, or clearly state that these metrics were computed only for the codeathon-internal analysis and are not yet available.
- [§5.1.4, §5.3, §5.4, §6] The Discussion claims that the codeathon 'successfully addressed several key challenges to working with metagenomic data at scale,' specifically '(b) rapidly producing benchmarks and reusable software.' The evidence in the manuscript is limited to early-stage infrastructure: Team 1 reports only a single positive control ('Found biosynthetic gene clusters in water sample with 56% similarity'), Team 3 presents a harness and pilot benchmark definitions but no tool evaluation results, and Team 4 describes a pipeline and a bottleneck (the 20-hour fasterq-dump step) without any quantitative comparison of the pipeline's sensitivity or speed against alternatives. These are perfectly reasonable proof-of-concept outcomes for a five-day codeathon, but the paper should either soften the 'benchmarks' language to 'pilot benchmark infrastructure' or provide the actual evaluation results for each team. As written, the Section 6 claim overstates the demonstrated accomplishments.
minor comments (7)
- [§1.2.1] The sentence 'Searching an entire database like SRA enables researches to explore multiple biological use-cases' contains a typo: 'researches' should be 'researchers.'
- [Table 2 and throughout] The software name is spelled inconsistently: 'MMseq2' appears in Table 2 and the text, while the canonical name is 'MMseqs2.' Please standardize.
- [§5.2.2] The sentence 'BLAST and MMseq2 agreed on a larger fraction (>141k containments)' uses the word 'fraction' for what appears to be a count; similarly, 'the Synteny-based method found fewer containments in common 75k' is ambiguous. Rephrase to state, for example, 'BLAST and MMseqs2 shared over 141,000 containments, while the synteny-based method shared 75,000 containments with the BLAST gold standard.'
- [Figure 7] Panel (b) is captioned 'Performance of the Snakemake pipeline obtained for BLAST and MMSeq2,' but it is not clear what 'performance' means (wall-clock time, memory, accuracy, or number of containments). Please add axis labels and a more descriptive caption.
- [§5.3.2] The phrase 'important aspect sof annotation' contains a typo: 'aspects.'
- [§6] The sentence 'A number of users spent significant amount of time and effort' is missing an article; it should read 'a significant amount of time.'
- [§5.2.1] The containment thresholds of 95% identity and 95% coverage are presented as definitions without justification. A brief sentence explaining the choice (e.g., relevance to strain-level containment or consistency with prior benchmarks) would help readers assess the design decision.
Circularity Check
Team 2's benchmark treats BLAST as its own gold standard, so the reported precision/recall comparisons are partially circular.
-
self definitional
[Section 5.2.2 (Datasets) and Section 5.2.3 (Codeathon product)]
"The results of using BLASTn for the search were treated as the gold standard. ... A ground truth set of containments was computed using BLAST. MMseq2 and Synteny-based methods were subsequently added to the Snakemake pipeline, allowing each tool's output to be evaluated against this standard."
BLAST is both the ground-truth generator and one of the tools being benchmarked: Table 2 lists BLAST as a tool whose pros/cons are evaluated, and Section 5.2.3 says precision/recall are computed for BLAST, Dashing, MiniMap2, and MashMap against the BLAST-defined standard. For BLAST, comparing its output to a truth set consisting of its own output makes its precision and recall definitionally perfect (or at least unable to reveal any BLAST-specific errors). For the other tools, every true positive, false positive, and false negative is defined by BLAST's sensitivity and false-positive profile, so the comparison cannot independently rank tools relative to biological truth. The reusable benchmark's load-bearing evaluation numbers reduce to BLAST's own output by construction.
full rationale
The paper is a codeathon report whose main deliverables are pipelines, repositories, and early-stage benchmark frameworks. Most of this infrastructure is independent and not circular: Team 1 uses BLAST as an external check on k-mer tools (BLAST is not being benchmarked there), Team 3 generates its own planted positives and uses standard sensitivity/FDR metrics, and Team 4 combines Pebblescout with ElasticBLAST without making one tool the arbiter of the other. The exception is Team 2's containment benchmark. Section 5.2.2 states repeatedly that 'The results of using BLASTn for the search were treated as the gold standard,' and Section 5.2.3 confirms that the ground-truth set 'was computed using BLAST' while BLAST itself is evaluated in the same pipeline (Table 2 and the stool metagenome precision/recall experiment). This means BLAST's own errors—both missed containments and spurious alignments—are silently encoded as truth, so the precision/recall values for BLAST and for every competing tool reduce to BLAST's output profile. That is a genuine partial circularity in the central benchmarking claim, though it does not invalidate the repositories, the pipelines, or the non-BLAST benchmark designs. I therefore assign a score of 6: one load-bearing evaluation reduces by construction, while the broader codeathon contribution retains substantial independent content.
Assumptions & free parameters
free parameters (4)
- Containment identity threshold =
95%
- Containment coverage threshold =
95%
- Default tool thresholds =
tool-specific defaults
- Pebblescout result limit =
2000 read sets
assumptions (4)
- domain assumption BLASTn output is a sufficient gold standard for DNA containment benchmarks
- domain assumption Containment definition with >95% identity and >95% coverage captures the relevant match notion
- domain assumption Pilot datasets (30 marine, 55 gut, 17 stool metagenomes) are sufficient to demonstrate benchmark pipelines
- domain assumption Cloud resources (NERSC, GCP, AWS) and tool versions are available as used in the codeathon
Cite this review
Pith. "Pith review of Contributions of the Petabyte Scale Sequence Search Codeathon toward efforts to scale sequence-based searches on SRA." pith.science (2026). https://pith.science/paper/RJYZZO7K
@misc{pith2026250506395,
author = {Pith},
title = {Pith review of: Contributions of the Petabyte Scale Sequence Search Codeathon toward efforts to scale sequence-based searches on SRA},
year = {2026},
howpublished = {\url{https://pith.science/paper/RJYZZO7K}},
note = {Machine review of arXiv:2505.06395}
}
read the original abstract
The volume of biological data being generated by the scientific community is growing exponentially, reflecting technological advances and research activities. The National Institutes of Health's (NIH) Sequence Read Archive (SRA), which is maintained by the National Center for Biotechnology Information (NCBI) at the National Library of Medicine (NLM), is a rapidly growing public database that researchers use to drive scientific discovery across all domains of life. This increase in available data has great promise for pushing scientific discovery but also introduces new challenges that scientific communities need to address. As genomic datasets have grown in scale and diversity, a parade of new methods and associated software have been developed to address the challenges posed by this growth. These methodological advances are vital for maximally leveraging the power of next-generation sequencing (NGS) technologies. With the goal of laying a foundation for evaluation of methods for petabyte-scale sequence search, the Department of Energy (DOE) Office of Biological and Environmental Research (BER), the NIH Office of Data Science Strategy (ODSS), and NCBI held a virtual codeathon 'Petabyte Scale Sequence Search: Metagenomics Benchmarking Codeathon' on September 27 - Oct 1 2021, to evaluate emerging solutions in petabyte scale sequence search. The codeathon attracted experts from national laboratories, research institutions, and universities across the world to (a) develop benchmarking approaches to address challenges in conducting large-scale analyses of metagenomic data (which comprises approximately 20% of SRA), (b) identify potential applications that benefit from SRA-wide searches and the tools required to execute the search, and (c) produce community resources i.e. a public facing repository with information to rebuild and reproduce the problems addressed by each team challenge.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
https://github.com/NCBI-Codeathons/psss-team2/wiki
-
[2]
https://www.ncbi.nlm.nih.gov/nuccore/NC_045512.2
-
[3]
https://github.com/NCBI-Codeathons/psss-team4/wiki
- [4]
- [5]
-
[6]
https://www.ncbi.nlm.nih.gov/nuccore/NC_002549.1
- [7]
- [8]
Show all 94 references
-
[9]
18 A PREPRINT
https://www.ncbi.nlm.nih.gov/nuccore/NC_024907.1. 18 A PREPRINT
-
[10]
Almodaresi, P
F. Almodaresi, P. Pandey, M. Ferdman, R. Johnson, and R. Patro. An efficient, scalable, and exact representation of high-dimensional color information enabled using de bruijn graph search. Journal of Computational Biology, 27(4):485–499, 2020
2020
-
[11]
S. F. Altschul, W. Gish, W. Miller, E. W. Myers, and D. J. Lipman. Basic local alignment search tool.Journal of molecular biology, 215(3):403–410, 1990
1990
-
[12]
S. F. Altschul, T. L. Madden, A. A. Schäffer, J. Zhang, Z. Zhang, W. Miller, and D. J. Lipman. Gapped blast and psi-blast: a new generation of protein database search programs. Nucleic acids research, 25(17):3389–3402, 1997
1997
-
[13]
Andreeva, D
A. Andreeva, D. Howorth, J.-M. Chandonia, S. E. Brenner, T. J. Hubbard, C. Chothia, and A. G. Murzin. Data growth and its impact on the scop database: new developments. Nucleic acids research, 36(suppl_1):D419–D425, 2007
2007
-
[14]
D. N. Baker and B. Langmead. Dashing: fast and accurate genomic distances with hyperloglog. Genome biology, 20:1–12, 2019
2019
-
[15]
Barrett, K
T. Barrett, K. Clark, R. Gevorgyan, V . Gorelenkov, E. Gribov, I. Karsch-Mizrachi, M. Kimelman, K. D. Pruitt, S. Resenchuk, T. Tatusova, et al. Bioproject and biosample databases at ncbi: facilitating capture and organization of metadata. Nucleic acids research, 40(D1):D57–D63, 2012
2012
-
[16]
D. A. Benson, K. Clark, I. Karsch-Mizrachi, D. J. Lipman, J. Ostell, and E. W. Sayers. Genbank. Nucleic acids research, 42(Database issue):D32, 2013
2013
-
[17]
Bingmann, P
T. Bingmann, P. Bradley, F. Gauger, and Z. Iqbal. Cobs: a compact bit-sliced signature index. InString Processing and Information Retrieval: 26th International Symposium, SPIRE 2019, Segovia, Spain, October 7–9, 2019, Proceedings 26, pages 285–303. Springer, 2019
2019
-
[18]
G. M. Boratyn, C. Camacho, P. S. Cooper, G. Coulouris, A. Fong, N. Ma, T. L. Madden, W. T. Matten, S. D. McGinnis, Y . Merezhuk, et al. Blast: a more efficient report with usability improvements. Nucleic acids research, 41(W1):W29–W33, 2013
2013
-
[19]
Bradley, H
P. Bradley, H. C. Den Bakker, E. P. Rocha, G. McVean, and Z. Iqbal. Ultrafast search of all deposited bacterial and viral genomic data. Nature biotechnology, 37(2):152–159, 2019
2019
-
[20]
Buchfink, K
B. Buchfink, K. Reuter, and H.-G. Drost. Sensitive protein alignments at tree-of-life scale using diamond. Nature methods, 18(4):366–368, 2021
2021
-
[21]
Camacho, G
C. Camacho, G. M. Boratyn, V . Joukov, R. Vera Alvarez, and T. L. Madden. Elasticblast: accelerating sequence search via cloud computing. BMC bioinformatics, 24(1):117, 2023
2023
-
[22]
Chikhi, B
R. Chikhi, B. Raffestin, A. Korobeynikov, R. Edgar, and A. Babaian. Logan: Planetary-scale genome assembly surveys life’s diversity. bioRxiv, 2024
2024
-
[23]
Di Tommaso, M
P. Di Tommaso, M. Chatzou, E. W. Floden, P. P. Barja, E. Palumbo, and C. Notredame. Nextflow enables reproducible computational workflows. Nature biotechnology, 35(4):316–319, 2017
2017
-
[24]
B. E. Dutilh, N. Cassman, K. McNair, S. E. Sanchez, G. G. Silva, L. Boling, J. J. Barr, D. R. Speth, V . Seguritan, R. K. Aziz, et al. A highly abundant bacteriophage discovered in the unknown sequences of human faecal metagenomes. Nature communications, 5(1):4498, 2014
2014
-
[25]
S. R. Eddy. Accelerated profile hmm searches. PLoS computational biology, 7(10):e1002195, 2011
2011
-
[26]
R. C. Edgar, J. Taylor, V . Lin, T. Altman, P. Barbera, D. Meleshko, D. Lohr, G. Novakovsky, B. Buchfink, B. Al-Shayeb, et al. Petabase-scale sequence alignment catalyses viral discovery. Nature, 602(7895):142–147, 2022
2022
-
[27]
El-Gebali, J
S. El-Gebali, J. Mistry, A. Bateman, S. R. Eddy, A. Luciani, S. C. Potter, M. Qureshi, L. J. Richardson, G. A. Salazar, A. Smart, et al. The pfam protein families database in 2019. Nucleic acids research, 47(D1):D427–D432, 2019
2019
-
[28]
R. D. Finn, J. Clements, and S. R. Eddy. Hmmer web server: interactive sequence similarity searching. Nucleic acids research, 39(suppl_2):W29–W37, 2011
2011
-
[29]
M. C. Frith, Y . Park, S. L. Sheetlin, and J. L. Spouge. The whole alignment and nothing but the alignment: the problem of spurious alignment flanks. Nucleic acids research, 36(18):5863–5871, 2008
2008
-
[30]
Glidden-Handgis and T
G. Glidden-Handgis and T. J. Wheeler. Was it a match i saw? approximate palindromes lead to overstated false match rates in benchmarks using reversed sequences. Bioinformatics Advances, 4(1):vbae052, 2024
2024
-
[31]
M. W. Gonzalez and W. R. Pearson. Homologous over-extension: a challenge for iterative similarity searches. Nucleic acids research, 38(7):2177–2189, 2010. 19 A PREPRINT
2010
-
[32]
Hoarfrost, N
A. Hoarfrost, N. Brown, C. T. Brown, and C. Arnosti. Sequencing data discovery with metaseek. Bioinformatics, 35(22):4857–4859, 2019
2019
-
[33]
Holley and P
G. Holley and P. Melsted. Bifrost: highly parallel construction and indexing of colored and compacted de bruijn graphs. Genome biology, 21:1–20, 2020
2020
-
[34]
Hubley, R
R. Hubley, R. D. Finn, J. Clements, S. R. Eddy, T. A. Jones, W. Bao, A. F. Smit, and T. J. Wheeler. The dfam database of repetitive dna families. Nucleic acids research, 44(D1):D81–D89, 2016
2016
-
[35]
NIH human microbiome project
Human Microbiome Project. NIH human microbiome project. https://www.hmpdacc.org/
-
[36]
Irber, N
L. Irber, N. T. Pierce-Ward, M. Abuelanin, H. Alexander, A. Anant, K. Barve, C. Baumler, O. Botvinnik, P. Brooks, D. Dsouza, et al. sourmash v4: A multitool to quickly search, compare, and analyze genomic and metagenomic data sets. Journal of Open Source Software, 9(98):6830, 2024
2024
-
[37]
Irber, N
L. Irber, N. T. Pierce-Ward, and C. T. Brown. Sourmash branchwater enables lightweight petabyte-scale sequence search. bioRxiv, pages 2022–11, 2022
2022
-
[38]
C. Jain, A. Dilthey, S. Koren, S. Aluru, and A. M. Phillippy. A fast approximate algorithm for mapping long reads to large reference databases. In International Conference on Research in Computational Molecular Biology, pages 66–81. Springer, 2017
2017
-
[39]
Karasikov, H
M. Karasikov, H. Mustafa, D. Danciu, C. Barber, M. Zimmermann, G. Rätsch, and A. Kahles. Metagraph: Indexing and analysing nucleotide archives at petabase-scale. BioRxiv, pages 2020–10, 2020
2020
-
[40]
Karplus, C
K. Karplus, C. Barrett, and R. Hughey. Hidden markov models for detecting remote protein homologies. Bioinformatics (Oxford, England), 14(10):846–856, 1998
1998
-
[41]
K. Katz, O. Shutov, R. Lapoint, M. Kimelman, J. R. Brister, and C. O’Sullivan. The sequence read archive: a decade more of explosive growth. Nucleic acids research, 50(D1):D387–D390, 2022
2022
-
[42]
K. S. Katz, O. Shutov, R. Lapoint, M. Kimelman, J. R. Brister, and C. O’Sullivan. Stat: a fast, scalable, minhash- based k-mer tool to assess sequence read archive next-generation sequence submissions. Genome Biology, 22(1):1–15, 2021
2021
-
[43]
S. M. Kiełbasa, R. Wan, K. Sato, P. Horton, and M. C. Frith. Adaptive seeds tame genomic sequence comparison. Genome research, 21(3):487–493, 2011
2011
-
[44]
Kodama, M
Y . Kodama, M. Shumway, and R. Leinonen. The sequence read archive: explosive growth of sequencing data. Nucleic acids research, 40(D1):D54–D56, 2012
2012
-
[45]
G. R. Krause, W. Shands, and T. J. Wheeler. Sensitive and error-tolerant annotation of protein-coding dna with bath. Bioinformatics Advances, page vbae088, 2024
2024
-
[46]
Krogh, M
A. Krogh, M. Brown, I. S. Mian, K. Sjölander, and D. Haussler. Hidden markov models in computational biology: Applications to protein modeling. Journal of molecular biology, 235(5):1501–1531, 1994
1994
-
[47]
T. W. Lab. Bagel: A framework for benchmarking sequence search tools. https://github.com/ TravisWheelerLab/BAGEL, 2024. Accessed: 2024-07-19
2024
-
[48]
A. A. Larkin, C. A. Garcia, N. Garcia, M. L. Brock, J. A. Lee, L. J. Ustick, L. Barbero, B. R. Carter, R. E. Sonnerup, L. D. Talley, et al. High spatial resolution global ocean metagenomes from bio-go-ship repeat hydrography transects. Scientific data, 8(1):107, 2021
2021
-
[49]
H. Li. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 34(18):3094–3100, 2018
2018
-
[50]
J. Li, H. Jia, X. Cai, H. Zhong, Q. Feng, S. Sunagawa, M. Arumugam, J. R. Kultima, E. Prifti, T. Nielsen, et al. An integrated catalog of reference genes in the human gut microbiome. Nature biotechnology, 32(8):834–841, 2014
2014
-
[51]
Liu and D
S. Liu and D. Koslicki. Cmash: fast, multi-resolution estimation of k-mer-based jaccard and containment indices. Bioinformatics, 38(Supplement_1):i28–i35, 2022
2022
-
[52]
B. Ma, M. T. France, J. Crabtree, J. B. Holm, M. S. Humphrys, R. M. Brotman, and J. Ravel. A comprehensive non-redundant gene catalog reveals extensive within-community intraspecies diversity in the human vagina. Nature communications, 11(1):940, 2020
2020
-
[53]
P. B. McGarvey, A. Nightingale, J. Luo, H. Huang, M. J. Martin, C. Wu, and U. Consortium. Uniprot genomic mapping for deciphering functional effects of missense variants. Human mutation, 40(6):694–705, 2019
2019
-
[54]
Mehringer, E
S. Mehringer, E. Seiler, F. Droop, M. Darvish, R. Rahn, M. Vingron, and K. Reinert. Hierarchical interleaved bloom filter: enabling ultrafast, approximate sequence queries. Genome Biology, 24(1):131, 2023
2023
-
[55]
Meyer, A
F. Meyer, A. Fritz, Z.-L. Deng, D. Koslicki, T. R. Lesker, A. Gurevich, G. Robertson, M. Alser, D. Antipov, F. Beghini, et al. Critical assessment of metagenome interpretation: the second round of challenges. Nature methods, 19(4):429–440, 2022. 20 A PREPRINT
2022
-
[56]
Mölder, K
F. Mölder, K. P. Jablonski, B. Letcher, M. B. Hall, C. H. Tomkins-Tinch, V . Sochat, J. Forster, S. Lee, S. O. Twardziok, A. Kanitz, et al. Sustainable data analysis with snakemake. F1000Research, 10, 2021
2021
-
[57]
About the STRIDES initiative
National Institutes of Health. About the STRIDES initiative. https://datascience.nih.gov/strides
-
[58]
SRA end user cloud access costs
National Institutes of Health. SRA end user cloud access costs. https://www.ncbi.nlm.nih.gov/sra/docs/ sra-cloud-access-costs/
-
[59]
Amazon web services joins NIH’s STRIDES initiative to harness latest cloud technologies for biomedical researchers, 2018
National Institutes of Health. Amazon web services joins NIH’s STRIDES initiative to harness latest cloud technologies for biomedical researchers, 2018. https://www.nih.gov/news-events/news-releases/ amazon-web-services-joins-nihs-strides-initiative-harness-latest-cloud-techno...
2018
-
[60]
NIH makes strides to accelerate discoveries in the cloud, 2018
National Institutes of Health. NIH makes strides to accelerate discoveries in the cloud, 2018. https://www.nih. gov/news-events/news-releases/nih-makes-strides-accelerate-discoveries-cloud
2018
-
[61]
NMDC metadata for soil study
National Microbiome Data Collaborative. NMDC metadata for soil study. https://data.microbiomedata. org/
-
[62]
NMDC workflows
National Microbiome Data Collaborative. NMDC workflows. https://nmdc-workflow-documentation. readthedocs.io/en/latest/chapters/overview.html#nmdc
-
[63]
E. P. Nawrocki, D. L. Kolbe, and S. R. Eddy. Infernal 1.0: inference of rna alignments. Bioinformatics, 25(10):1335–1337, 2009
2009
-
[64]
H. B. Nielsen, M. Almeida, A. S. Juncker, S. Rasmussen, J. Li, S. Sunagawa, D. R. Plichta, L. Gautier, A. G. Pedersen, E. Le Chatelier, et al. Identification and assembly of genomes and genetic elements in complex metagenomic samples without using reference genomes. Nature bio...
2014
-
[65]
S. Nurk, S. Koren, A. Rhie, M. Rautiainen, A. V . Bzikadze, A. Mikheenko, M. R. V ollger, N. Altemose, L. Uralsky, A. Gershman, et al. The complete sequence of a human genome. Science, 376(6588):44–53, 2022
2022
-
[66]
B. D. Ondov, G. J. Starrett, A. Sappington, A. Kostic, S. Koren, C. B. Buck, and A. M. Phillippy. Mash screen: high-throughput sequence containment estimation for genome discovery. Genome biology, 20:1–13, 2019
2019
-
[67]
B. D. Ondov, T. J. Treangen, P. Melsted, A. B. Mallonee, N. H. Bergman, S. Koren, and A. M. Phillippy. Mash: fast genome and metagenome distance estimation using minhash. Genome biology, 17:1–14, 2016
2016
-
[68]
Pandey, F
P. Pandey, F. Almodaresi, M. A. Bender, M. Ferdman, R. Johnson, and R. Patro. Mantis: a fast, small, and exact large-scale sequence-search index. Cell systems, 7(2):201–207, 2018
2018
-
[69]
G. W. Park, T. F. F. Ng, A. L. Freeland, V . C. Marconi, J. A. Boom, M. A. Staat, A. M. Montmayeur, H. Browne, J. Narayanan, D. C. Payne, et al. Crassphage as a novel tool to detect human fecal contamination on environmental surfaces and hands. Emerging Infectious Diseases, 26...
2020
-
[70]
D. H. Parks, C. Rinke, M. Chuvochina, P.-A. Chaumeil, B. J. Woodcroft, P. N. Evans, P. Hugenholtz, and G. W. Tyson. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nature microbiology, 2(11):1533–1542, 2017
2017
-
[71]
N. T. Pierce, L. Irber, T. Reiter, P. Brooks, and C. T. Brown. Large-scale sequence comparisons with sourmash. F1000Research, 8, 2019
2019
-
[72]
Pruitt, K
K. Pruitt, K. Clark, T. Tatusova, and I. Mizrachi. Bioproject help. In BioProject Help [Internet]. National Center for Biotechnology Information (US), 2011
2011
-
[73]
J. Qin, R. Li, J. Raes, M. Arumugam, K. S. Burgdorf, C. Manichanh, T. Nielsen, N. Pons, F. Levenez, T. Yamada, et al. A human gut microbial gene catalogue established by metagenomic sequencing. nature, 464(7285):59–65, 2010
2010
-
[74]
J. W. Roddy, D. H. Rich, and T. J. Wheeler. nail: software for high-speed, high-sensitivity protein sequence annotation. bioRxiv, 2024
2024
-
[75]
Seiler, S
E. Seiler, S. Mehringer, M. Darvish, E. Turc, and K. Reinert. Raptor: A fast and space-efficient pre-filter for querying very large collections of nucleotide sequences. Iscience, 24(7), 2021
2021
-
[76]
Sharon, M
I. Sharon, M. J. Morowitz, B. C. Thomas, E. K. Costello, D. A. Relman, and J. F. Banfield. Time series community genomics analysis reveals rapid shifts in bacterial species, strains, and phage during infant gut colonization. Genome research, 23(1):111–120, 2013
2013
-
[77]
S. A. Shiryev and R. Agarwala. Indexing and searching petabase-scale nucleotide resources. Nature methods, 21(6):994–1002, 2024
2024
-
[78]
Shumway, G
M. Shumway, G. Cochrane, and H. Sugawara. Archiving next generation sequencing data. Nucleic acids research, 38(suppl_1):D870–D871, 2010. 21 A PREPRINT
2010
-
[79]
J. Söding. Protein homology detection by hmm–hmm comparison. Bioinformatics, 21(7):951–960, 2005
2005
-
[80]
Solomon and C
B. Solomon and C. Kingsford. Fast search of thousands of short-read sequencing experiments. Nature biotechnol- ogy, 34(3):300–302, 2016
2016
-
[81]
branchwater: Searching large collections of sequencing data with genome-scale queries
sourmash. branchwater: Searching large collections of sequencing data with genome-scale queries. https: //github.com/sourmash-bio/branchwater
-
[82]
Steinegger and J
M. Steinegger and J. Söding. Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature biotechnology, 35(11):1026–1028, 2017
2017
-
[83]
B. J. Tully, E. D. Graham, and J. F. Heidelberg. The reconstruction of 2,631 draft metagenome-assembled genomes from the global oceans. Scientific data, 5(1):1–8, 2018
2018
-
[84]
P. J. Turnbaugh, R. E. Ley, M. Hamady, C. M. Fraser-Liggett, R. Knight, and J. I. Gordon. The human microbiome project. Nature, 449(7164):804–810, 2007
2007
-
[85]
J. C. Venter, M. D. Adams, E. W. Myers, P. W. Li, R. J. Mural, G. G. Sutton, H. O. Smith, M. Yandell, C. A. Evans, R. A. Holt, et al. The sequence of the human genome. science, 291(5507):1304–1351, 2001
2001
-
[86]
T. J. Wheeler, J. Clements, S. R. Eddy, R. Hubley, T. A. Jones, J. Jurka, A. F. Smit, and R. D. Finn. Dfam: a database of repetitive dna based on profile hidden markov models. Nucleic acids research, 41(D1):D70–D82, 2012
2012
-
[87]
T. J. Wheeler and S. R. Eddy. nhmmer: Dna homology search with profile hmms. Bioinformatics, 29(19):2487– 2489, 2013
2013
-
[88]
M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, et al. The FAIR guiding principles for scientific data management and stewardship. Scientific data, 3(1):1–9, 2016
2016
-
[89]
D. E. Wood, J. Lu, and B. Langmead. Improved metagenomic analysis with kraken 2. Genome biology, 20:1–13, 2019
2019
-
[90]
D. E. Wood and S. L. Salzberg. Kraken: ultrafast metagenomic sequence classification using exact alignments. Genome biology, 15:1–12, 2014
2014
-
[91]
E. S. Wright. Using decipher v2. 0 to analyze big biological sequence data in r. R Journal, 8(1), 2016
2016
-
[92]
Wu and Y
S. Wu and Y . Zhang. Muster: improving protein sequence profile–profile alignments by using multiple sources of structure information. Proteins: Structure, Function, and Bioinformatics, 72(2):547–556, 2008
2008
-
[93]
J. Ye, S. McGinnis, and T. L. Madden. Blast: improvements for better sequence analysis. Nucleic acids research, 34(suppl_2):W6–W9, 2006
2006
-
[94]
Y . Yu, J. Liu, X. Liu, Y . Zhang, E. Magner, E. Lehnert, C. Qian, and J. Liu. Seqothello: querying rna-seq experiments at scale. Genome biology, 19:1–13, 2018. 7 Appendix 7.1 Appendix 1: Marine metagenome data accessions Ocean metagenome dataset used to generate gold standard...
2018
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.