Pith. sign in

REVIEW 3 major objections 7 minor 94 references

Contributions of the Petabyte Scale Sequence Search Codeathon toward efforts to scale sequence-based searches on SRA

T0 review · 3 major / 7 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Four codeathon teams produced reusable benchmarks and proof-of-concept pipelines for petabyte-scale sequence search, the paper reports.

desk verdict An honestly scoped codeathon report with real, reusable artifacts; the BLAST-as-gold-standard circularity is genuine but does not sink its central claim. read the letter →

arxiv 2505.06395 v1 pith:RJYZZO7K submitted 2025-05-09 q-bio.OT

classification q-bio.OT
keywords metagenomicssequencesearchpetabytescalebenchmarkingSRAk-merindexingalignmentcodeathon
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports that a focused codeathon produced four working proof-of-concept projects aimed at benchmarking sequence search against a petabyte-scale public sequence archive. Team 1 built a pipeline to screen metagenomes for user-provided long queries with k-mer search tools; Team 2 generated gold-standard datasets and an evaluator for contig containment; Team 3 created hard-annotation benchmarks and a harness to run them across computing platforms; Team 4 connected a k-mer index to a cloud alignment engine, showing a search-to-alignment workflow. The central claim is that these early-stage benchmarks, pipelines, datasets, and public repositories form a reusable foundation for evaluating tools that could eventually let researchers search the whole archive by sequence content.

What carries the argument

The machinery is a set of benchmark blueprints: a gold-standard definition of sequence containment; the use of BLASTn as ground truth; precision/recall evaluation against that ground truth; workflow layers built with Snakemake and Nextflow; the BAGEL harness for hard-annotation benchmarks; and the two-stage Pebblescout-to-ElasticBLAST search. The containment definition and BLAST ground truth carry the evaluation numbers, the workflow frameworks carry reproducibility, and the two-stage pipeline carries the demonstration that an index can narrow a petabyte-scale corpus to a few hundred plausible candidates before alignment.

What would settle it

Take a metagenome sample, spike in contigs at known identities and lengths, then run the gold-standard construction and count how many spiked containments BLASTn misses and how many unrelated pairs it reports; if BLASTn disagrees with the spike-in truth by more than a small margin, the precision and recall rankings in the benchmarks do not reflect true search accuracy.

Watch

Extended reading notes

Core claim

The codeathon's contribution is a decomposition of petabyte-scale sequence search into benchmarkable subproblems, each supported by a working prototype. Team 1's bothie workflow wraps several k-mer sketching tools and uses BLAST-based comparisons as likely ground truth to detect long query sequences inside metagenomes. Team 2 defined containment as an alignment with more than 95% identity covering more than 95% of the shorter contig, generated gold standards by running BLASTn, and built a Snakemake pipeline that scores other tools by precision and recall stratified by identity and contig length. Team 3 produced pilot hard-annotation benchmarks and the BAGEL harness for running them on local, cluster, and cloud environments. Team 4 combined the Pebblescout k-mer index with ElasticBLAST to prefilter candidate SRA runs and then align queries against them, finding that the main bottleneck was downloading the selected reads rather than the search itself. Together these proofs-of-concept, with the datasets and repositories shared alongside them, are the paper's contribution.

Load-bearing premise

The benchmarking results assume that BLASTn's alignments are a correct and complete gold standard for DNA containment, so any match BLAST misses or falsely reports is inherited by every tool's precision and recall numbers.

Editorial extensions

If this is right

  • Researchers can reproduce and extend the codeathon benchmarks from public repositories rather than reimplementing evaluation from scratch.
  • A biologist with a query genome or biosynthetic gene cluster can use the bothie pipeline to screen metagenomes for near-identical containing samples.
  • The Team 2 gold standards provide a baseline against which new containment tools can be scored by precision and recall, stratified by identity and contig length.
  • The Team 4 pipeline shows that after a k-mer index reduces the archive to candidate runs, the practical limit becomes download time rather than search time.
  • The BAGEL harness lets tool developers add a new alignment or annotation tool by wrapping it in a container, without writing cluster or cloud code.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: because the pipelines and repositories are public, a corrected gold standard could be swapped in without rebuilding the evaluation framework, so the infrastructure's value does not depend on BLAST being perfect.
  • Inference: the narrow containment definition (more than 95% identity over more than 95% of the shorter contig) excludes partial or diverged matches; rerunning the benchmarks with a looser definition may change which tools look best for viral-discovery use cases.
  • Inference: Team 2's proposed order-preserving minimizer sketch, if completed into a colored de Bruijn graph, could make containment queries nearly constant-time and is a natural next step to test within the same harness.
  • Inference: the download bottleneck identified by Team 4 suggests that cloud-native formats and direct streaming from object storage, rather than faster alignment, are the near-term limiting factor for SRA-wide pipelines.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The manuscript reports on a virtual codeathon held September 27 - October 1, 2021, convened by DOE BER, NIH ODSS, and NCBI to develop benchmarking approaches and reusable software for petabyte-scale sequence-based search of the NCBI Sequence Read Archive, with a focus on metagenomic data. Four teams produced early-stage deliverables: Team 1 developed 'bothie,' a pipeline for detecting user-provided long query sequences with k-mer-based tools using BLAST as a 'likely groundtruth'; Team 2 developed a containment-benchmark pipeline with BLASTn as the gold standard and evaluated MMseqs2, Minimap2, dashing, MashMap, and a synteny-based method; Team 3 developed the BAGEL harness and pilot hard-annotation benchmarks, and Team 4 built a combined Pebblescout + ElasticBLAST pipeline. The Discussion claims the codeathon successfully produced benchmarks, reusable software, and a cloud-based foundation for scaling SRA searches. The paper is written as a codeathon report with modest, explicitly preliminary claims, and it emphasizes reproducibility through public repositories and workflow tools.

Significance. The paper's strengths are its documentation of a structured community effort, the public availability of the four team repositories (bothie, psss-team2, psss-team3-hard-annotation, psss-team4), and the candid description of practical bottlenecks such as the fasterq-dump transfer problem. As an organizational and infrastructure contribution, the manuscript has value for future codeathon organizers and for researchers seeking reusable scaffolding for metagenomic search workflows. However, the central claim of producing 'benchmarks' is not backed by reported evaluation results: no precision/recall numbers, error bars, or statistical analyses appear in the text. The most load-bearing assumption, BLAST as gold standard while BLAST itself is a benchmarked tool, makes the Team 2 evaluation circular and limits the value of the claimed benchmark as an independent resource. The paper would be strengthened by softening the 'benchmark' framing or by presenting actual metrics and an accompanying validation strategy.

major comments (3)
  1. [§5.2.2, Table 2] The benchmarking methodology in Section 5.2.2 treats BLASTn output as the gold standard for containment ('The results of using BLASTn for the search were treated as the gold standard'), while BLAST is itself one of the tools evaluated in Table 2. This makes the precision/recall calculations circular: every true positive, false positive, and false negative for MMseqs2, Minimap2, dashing, MashMap, and synteny is defined relative to BLAST's own sensitivity and false-positive profile. A containment missed by BLAST is scored as a false negative for the other tools, and a spurious BLAST hit is scored as ground truth. This undermines the claim in Section 6 that the codeathon produced 'benchmarks and reusable software' as an evaluation resource, because the benchmark cannot independently rank methods whose errors are defined by the very tool being assessed. I recommend either (i) constructing ground truth from simulated reads/contigs with known containment, (ii) using a consensus of several independent aligners, or (iii) clearly relabeling all reported precision/recall as 'agreement with BLAST' rather than as absolute accuracy.
  2. [§5.2.2 and §5.2.3] The text states that precision and recall were computed for the stool metagenome dataset ('Containments were identified using all vs all alignments to obtain precision and recall values on BLAST, Dashing, MiniMap2 and MashMap'), and Section 5.2.3 describes TP/FP/FN classification and stratification by percent identity, contig length, and confidence. However, no precision, recall, TP/FP/FN, or stratified values are reported anywhere in the paper. The only quantitative results are aggregate containment counts in Section 5.2.2 (155,372 total containments; >141k BLAST/MMseq2 agreement; 75k synteny overlap). Since the central deliverable is a benchmark, the absence of the actual benchmark scores makes the results non-reproducible as stated and weakens the Section 6 claim of 'rapidly producing benchmarks.' Please add tables or figures with the precision/recall and runtime/memory measurements mentioned in Section 5.2.3, or clearly state that these metrics were computed only for the codeathon-internal analysis and are not yet available.
  3. [§5.1.4, §5.3, §5.4, §6] The Discussion claims that the codeathon 'successfully addressed several key challenges to working with metagenomic data at scale,' specifically '(b) rapidly producing benchmarks and reusable software.' The evidence in the manuscript is limited to early-stage infrastructure: Team 1 reports only a single positive control ('Found biosynthetic gene clusters in water sample with 56% similarity'), Team 3 presents a harness and pilot benchmark definitions but no tool evaluation results, and Team 4 describes a pipeline and a bottleneck (the 20-hour fasterq-dump step) without any quantitative comparison of the pipeline's sensitivity or speed against alternatives. These are perfectly reasonable proof-of-concept outcomes for a five-day codeathon, but the paper should either soften the 'benchmarks' language to 'pilot benchmark infrastructure' or provide the actual evaluation results for each team. As written, the Section 6 claim overstates the demonstrated accomplishments.
minor comments (7)
  1. [§1.2.1] The sentence 'Searching an entire database like SRA enables researches to explore multiple biological use-cases' contains a typo: 'researches' should be 'researchers.'
  2. [Table 2 and throughout] The software name is spelled inconsistently: 'MMseq2' appears in Table 2 and the text, while the canonical name is 'MMseqs2.' Please standardize.
  3. [§5.2.2] The sentence 'BLAST and MMseq2 agreed on a larger fraction (>141k containments)' uses the word 'fraction' for what appears to be a count; similarly, 'the Synteny-based method found fewer containments in common 75k' is ambiguous. Rephrase to state, for example, 'BLAST and MMseqs2 shared over 141,000 containments, while the synteny-based method shared 75,000 containments with the BLAST gold standard.'
  4. [Figure 7] Panel (b) is captioned 'Performance of the Snakemake pipeline obtained for BLAST and MMSeq2,' but it is not clear what 'performance' means (wall-clock time, memory, accuracy, or number of containments). Please add axis labels and a more descriptive caption.
  5. [§5.3.2] The phrase 'important aspect sof annotation' contains a typo: 'aspects.'
  6. [§6] The sentence 'A number of users spent significant amount of time and effort' is missing an article; it should read 'a significant amount of time.'
  7. [§5.2.1] The containment thresholds of 95% identity and 95% coverage are presented as definitions without justification. A brief sentence explaining the choice (e.g., relevance to strain-level containment or consistency with prior benchmarks) would help readers assess the design decision.

Circularity Check

1 steps flagged · score 6.0 of 10

Team 2's benchmark treats BLAST as its own gold standard, so the reported precision/recall comparisons are partially circular.

  1. self definitional [Section 5.2.2 (Datasets) and Section 5.2.3 (Codeathon product)]
    "The results of using BLASTn for the search were treated as the gold standard. ... A ground truth set of containments was computed using BLAST. MMseq2 and Synteny-based methods were subsequently added to the Snakemake pipeline, allowing each tool's output to be evaluated against this standard."

    BLAST is both the ground-truth generator and one of the tools being benchmarked: Table 2 lists BLAST as a tool whose pros/cons are evaluated, and Section 5.2.3 says precision/recall are computed for BLAST, Dashing, MiniMap2, and MashMap against the BLAST-defined standard. For BLAST, comparing its output to a truth set consisting of its own output makes its precision and recall definitionally perfect (or at least unable to reveal any BLAST-specific errors). For the other tools, every true positive, false positive, and false negative is defined by BLAST's sensitivity and false-positive profile, so the comparison cannot independently rank tools relative to biological truth. The reusable benchmark's load-bearing evaluation numbers reduce to BLAST's own output by construction.

full rationale

The paper is a codeathon report whose main deliverables are pipelines, repositories, and early-stage benchmark frameworks. Most of this infrastructure is independent and not circular: Team 1 uses BLAST as an external check on k-mer tools (BLAST is not being benchmarked there), Team 3 generates its own planted positives and uses standard sensitivity/FDR metrics, and Team 4 combines Pebblescout with ElasticBLAST without making one tool the arbiter of the other. The exception is Team 2's containment benchmark. Section 5.2.2 states repeatedly that 'The results of using BLASTn for the search were treated as the gold standard,' and Section 5.2.3 confirms that the ground-truth set 'was computed using BLAST' while BLAST itself is evaluated in the same pipeline (Table 2 and the stool metagenome precision/recall experiment). This means BLAST's own errors—both missed containments and spurious alignments—are silently encoded as truth, so the precision/recall values for BLAST and for every competing tool reduce to BLAST's output profile. That is a genuine partial circularity in the central benchmarking claim, though it does not invalidate the repositories, the pipelines, or the non-BLAST benchmark designs. I therefore assign a score of 6: one load-bearing evaluation reduces by construction, while the broader codeathon contribution retains substantial independent content.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The codeathon results depend on a small number of hand-chosen thresholds (95% identity/coverage, tool defaults, Pebblescout 2,000-read limit) and on the domain assumption that BLAST is a valid gold standard. No novel theoretical entities are introduced; the new software artifacts (bothie, BAGEL, team pipelines) are concrete deliverables with public repositories.

free parameters (4)
  • Containment identity threshold = 95%
    Section 5.2.1: 'higher than 95% identity within the aligned region'. Chosen by hand, not derived; defines truth in Team 2 benchmarks.
  • Containment coverage threshold = 95%
    Section 5.2.1: 'covers more than 95% of the length of the shorter contig'. Chosen by hand; controls which contig pairs count as matches.
  • Default tool thresholds = tool-specific defaults
    Section 5.1.3: 'detection was defined as containment-based similarity above the default threshold of each tool'. These vary by tool and are not uniformly justified.
  • Pebblescout result limit = 2000 read sets
    Section 5.4.2: 'returning search results to a maximum of 2000 read sets (at the time) in the output by default'. Constrains Team 4 output.
assumptions (4)
  • domain assumption BLASTn output is a sufficient gold standard for DNA containment benchmarks
    Section 5.2.2 states 'The results of using BLASTn for the search were treated as the gold standard.' This underpins all Team 2 precision/recall values and Team 1 groundtruth comparisons.
  • domain assumption Containment definition with >95% identity and >95% coverage captures the relevant match notion
    Section 5.2.1 defines containment with these thresholds; no independent biological or statistical basis is given for the cutoffs.
  • domain assumption Pilot datasets (30 marine, 55 gut, 17 stool metagenomes) are sufficient to demonstrate benchmark pipelines
    Sections 5.2.2 and 5.2.4 derive containment counts and performance observations from these small sets; the paper frames them as proof-of-concept.
  • domain assumption Cloud resources (NERSC, GCP, AWS) and tool versions are available as used in the codeathon
    All teams relied on cloud infrastructure; exact tool versions and cloud configurations are not all pinned, limiting exact reproduction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Contributions of the Petabyte Scale Sequence Search Codeathon toward efforts to scale sequence-based searches on SRA." pith.science (2026). https://pith.science/paper/RJYZZO7K

@misc{pith2026250506395,
  author       = {Pith},
  title        = {Pith review of: Contributions of the Petabyte Scale Sequence Search Codeathon toward efforts to scale sequence-based searches on SRA},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RJYZZO7K}},
  note         = {Machine review of arXiv:2505.06395}
}
read the original abstract

The volume of biological data being generated by the scientific community is growing exponentially, reflecting technological advances and research activities. The National Institutes of Health's (NIH) Sequence Read Archive (SRA), which is maintained by the National Center for Biotechnology Information (NCBI) at the National Library of Medicine (NLM), is a rapidly growing public database that researchers use to drive scientific discovery across all domains of life. This increase in available data has great promise for pushing scientific discovery but also introduces new challenges that scientific communities need to address. As genomic datasets have grown in scale and diversity, a parade of new methods and associated software have been developed to address the challenges posed by this growth. These methodological advances are vital for maximally leveraging the power of next-generation sequencing (NGS) technologies. With the goal of laying a foundation for evaluation of methods for petabyte-scale sequence search, the Department of Energy (DOE) Office of Biological and Environmental Research (BER), the NIH Office of Data Science Strategy (ODSS), and NCBI held a virtual codeathon 'Petabyte Scale Sequence Search: Metagenomics Benchmarking Codeathon' on September 27 - Oct 1 2021, to evaluate emerging solutions in petabyte scale sequence search. The codeathon attracted experts from national laboratories, research institutions, and universities across the world to (a) develop benchmarking approaches to address challenges in conducting large-scale analyses of metagenomic data (which comprises approximately 20% of SRA), (b) identify potential applications that benefit from SRA-wide searches and the tools required to execute the search, and (c) produce community resources i.e. a public facing repository with information to rebuild and reproduce the problems addressed by each team challenge.

Figures

Figures reproduced from arXiv: 2505.06395 by the authors.

Figure 1
Figure 1. Growth of SRA. Databases like SRA are growing rapidly and are used extensively by scientific communities. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. SRA cumulative records growth over time. ‘runs’ refer to SRA accessions and ‘samples’ correspond to [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Percent SRA accessions by organism as of December 2024. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Top metagenomic accession record counts classified based on type ‘scientific organism name’ as of Dec 2024. [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Codeathon planning overview. NIH 12% DOE National Laboratories 24% Industry 12% Academia 52% [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Codeathon Team Projects overview. 5 Team Projects 5.1 Team 1: Identifying metagenomic samples with user-provided long queries 5.1.1 Background Microbial environments are being explored on a genomic level like never before, with hundreds of thousands of metagenome datas…
Figure 7
Figure 7. Figure 7: Preliminary results obtained for the marine metagenome dataset [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Proposed workflow [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Proposed Snakemake workflow • Acidianus filamentous virus 1: NC_005830.1 [8] • Ancient caribou feces associated virus: NC_024907.1 [9] 5.4.3 Codeathon product Major steps of the pipeline and corresponding Snakemake workflow (as illustrated in Figs. 8 and 9 respectively…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

94 extracted references · 79 canonical work pages

  1. [1]

    https://github.com/NCBI-Codeathons/psss-team2/wiki

  2. [2]

    https://www.ncbi.nlm.nih.gov/nuccore/NC_045512.2

  3. [3]

    https://github.com/NCBI-Codeathons/psss-team4/wiki

  4. [4]

    https://www.ncbi.nlm.nih.gov/nuccore/29366675?sat=50&satkey=26787819

  5. [5]

    https://www.ncbi.nlm.nih.gov/nuccore/556503834?sat=50&satkey=48187781

  6. [6]

    https://www.ncbi.nlm.nih.gov/nuccore/NC_002549.1

  7. [7]

    https://www.ncbi.nlm.nih.gov/nuccore/194100415?sat=50&satkey=26790357

  8. [8]

    https://www.ncbi.nlm.nih.gov/nuccore/45655866?sat=51&satkey=72306780

Show all 94 references
  1. [9]

    18 A PREPRINT

    https://www.ncbi.nlm.nih.gov/nuccore/NC_024907.1. 18 A PREPRINT

  2. [10]

    Almodaresi, P

    F. Almodaresi, P. Pandey, M. Ferdman, R. Johnson, and R. Patro. An efficient, scalable, and exact representation of high-dimensional color information enabled using de bruijn graph search. Journal of Computational Biology, 27(4):485–499, 2020

  3. [11]

    S. F. Altschul, W. Gish, W. Miller, E. W. Myers, and D. J. Lipman. Basic local alignment search tool.Journal of molecular biology, 215(3):403–410, 1990

  4. [12]

    S. F. Altschul, T. L. Madden, A. A. Schäffer, J. Zhang, Z. Zhang, W. Miller, and D. J. Lipman. Gapped blast and psi-blast: a new generation of protein database search programs. Nucleic acids research, 25(17):3389–3402, 1997

  5. [13]

    Andreeva, D

    A. Andreeva, D. Howorth, J.-M. Chandonia, S. E. Brenner, T. J. Hubbard, C. Chothia, and A. G. Murzin. Data growth and its impact on the scop database: new developments. Nucleic acids research, 36(suppl_1):D419–D425, 2007

  6. [14]

    D. N. Baker and B. Langmead. Dashing: fast and accurate genomic distances with hyperloglog. Genome biology, 20:1–12, 2019

  7. [15]

    Barrett, K

    T. Barrett, K. Clark, R. Gevorgyan, V . Gorelenkov, E. Gribov, I. Karsch-Mizrachi, M. Kimelman, K. D. Pruitt, S. Resenchuk, T. Tatusova, et al. Bioproject and biosample databases at ncbi: facilitating capture and organization of metadata. Nucleic acids research, 40(D1):D57–D63, 2012

  8. [16]

    D. A. Benson, K. Clark, I. Karsch-Mizrachi, D. J. Lipman, J. Ostell, and E. W. Sayers. Genbank. Nucleic acids research, 42(Database issue):D32, 2013

  9. [17]

    Bingmann, P

    T. Bingmann, P. Bradley, F. Gauger, and Z. Iqbal. Cobs: a compact bit-sliced signature index. InString Processing and Information Retrieval: 26th International Symposium, SPIRE 2019, Segovia, Spain, October 7–9, 2019, Proceedings 26, pages 285–303. Springer, 2019

  10. [18]

    G. M. Boratyn, C. Camacho, P. S. Cooper, G. Coulouris, A. Fong, N. Ma, T. L. Madden, W. T. Matten, S. D. McGinnis, Y . Merezhuk, et al. Blast: a more efficient report with usability improvements. Nucleic acids research, 41(W1):W29–W33, 2013

  11. [19]

    Bradley, H

    P. Bradley, H. C. Den Bakker, E. P. Rocha, G. McVean, and Z. Iqbal. Ultrafast search of all deposited bacterial and viral genomic data. Nature biotechnology, 37(2):152–159, 2019

  12. [20]

    Buchfink, K

    B. Buchfink, K. Reuter, and H.-G. Drost. Sensitive protein alignments at tree-of-life scale using diamond. Nature methods, 18(4):366–368, 2021

  13. [21]

    Camacho, G

    C. Camacho, G. M. Boratyn, V . Joukov, R. Vera Alvarez, and T. L. Madden. Elasticblast: accelerating sequence search via cloud computing. BMC bioinformatics, 24(1):117, 2023

  14. [22]

    Chikhi, B

    R. Chikhi, B. Raffestin, A. Korobeynikov, R. Edgar, and A. Babaian. Logan: Planetary-scale genome assembly surveys life’s diversity. bioRxiv, 2024

  15. [23]

    Di Tommaso, M

    P. Di Tommaso, M. Chatzou, E. W. Floden, P. P. Barja, E. Palumbo, and C. Notredame. Nextflow enables reproducible computational workflows. Nature biotechnology, 35(4):316–319, 2017

  16. [24]

    B. E. Dutilh, N. Cassman, K. McNair, S. E. Sanchez, G. G. Silva, L. Boling, J. J. Barr, D. R. Speth, V . Seguritan, R. K. Aziz, et al. A highly abundant bacteriophage discovered in the unknown sequences of human faecal metagenomes. Nature communications, 5(1):4498, 2014

  17. [25]

    S. R. Eddy. Accelerated profile hmm searches. PLoS computational biology, 7(10):e1002195, 2011

  18. [26]

    R. C. Edgar, J. Taylor, V . Lin, T. Altman, P. Barbera, D. Meleshko, D. Lohr, G. Novakovsky, B. Buchfink, B. Al-Shayeb, et al. Petabase-scale sequence alignment catalyses viral discovery. Nature, 602(7895):142–147, 2022

  19. [27]

    El-Gebali, J

    S. El-Gebali, J. Mistry, A. Bateman, S. R. Eddy, A. Luciani, S. C. Potter, M. Qureshi, L. J. Richardson, G. A. Salazar, A. Smart, et al. The pfam protein families database in 2019. Nucleic acids research, 47(D1):D427–D432, 2019

  20. [28]

    R. D. Finn, J. Clements, and S. R. Eddy. Hmmer web server: interactive sequence similarity searching. Nucleic acids research, 39(suppl_2):W29–W37, 2011

  21. [29]

    M. C. Frith, Y . Park, S. L. Sheetlin, and J. L. Spouge. The whole alignment and nothing but the alignment: the problem of spurious alignment flanks. Nucleic acids research, 36(18):5863–5871, 2008

  22. [30]

    Glidden-Handgis and T

    G. Glidden-Handgis and T. J. Wheeler. Was it a match i saw? approximate palindromes lead to overstated false match rates in benchmarks using reversed sequences. Bioinformatics Advances, 4(1):vbae052, 2024

  23. [31]

    M. W. Gonzalez and W. R. Pearson. Homologous over-extension: a challenge for iterative similarity searches. Nucleic acids research, 38(7):2177–2189, 2010. 19 A PREPRINT

  24. [32]

    Hoarfrost, N

    A. Hoarfrost, N. Brown, C. T. Brown, and C. Arnosti. Sequencing data discovery with metaseek. Bioinformatics, 35(22):4857–4859, 2019

  25. [33]

    Holley and P

    G. Holley and P. Melsted. Bifrost: highly parallel construction and indexing of colored and compacted de bruijn graphs. Genome biology, 21:1–20, 2020

  26. [34]

    Hubley, R

    R. Hubley, R. D. Finn, J. Clements, S. R. Eddy, T. A. Jones, W. Bao, A. F. Smit, and T. J. Wheeler. The dfam database of repetitive dna families. Nucleic acids research, 44(D1):D81–D89, 2016

  27. [35]

    NIH human microbiome project

    Human Microbiome Project. NIH human microbiome project. https://www.hmpdacc.org/

  28. [36]

    Irber, N

    L. Irber, N. T. Pierce-Ward, M. Abuelanin, H. Alexander, A. Anant, K. Barve, C. Baumler, O. Botvinnik, P. Brooks, D. Dsouza, et al. sourmash v4: A multitool to quickly search, compare, and analyze genomic and metagenomic data sets. Journal of Open Source Software, 9(98):6830, 2024

  29. [37]

    Irber, N

    L. Irber, N. T. Pierce-Ward, and C. T. Brown. Sourmash branchwater enables lightweight petabyte-scale sequence search. bioRxiv, pages 2022–11, 2022

  30. [38]

    C. Jain, A. Dilthey, S. Koren, S. Aluru, and A. M. Phillippy. A fast approximate algorithm for mapping long reads to large reference databases. In International Conference on Research in Computational Molecular Biology, pages 66–81. Springer, 2017

  31. [39]

    Karasikov, H

    M. Karasikov, H. Mustafa, D. Danciu, C. Barber, M. Zimmermann, G. Rätsch, and A. Kahles. Metagraph: Indexing and analysing nucleotide archives at petabase-scale. BioRxiv, pages 2020–10, 2020

  32. [40]

    Karplus, C

    K. Karplus, C. Barrett, and R. Hughey. Hidden markov models for detecting remote protein homologies. Bioinformatics (Oxford, England), 14(10):846–856, 1998

  33. [41]

    K. Katz, O. Shutov, R. Lapoint, M. Kimelman, J. R. Brister, and C. O’Sullivan. The sequence read archive: a decade more of explosive growth. Nucleic acids research, 50(D1):D387–D390, 2022

  34. [42]

    K. S. Katz, O. Shutov, R. Lapoint, M. Kimelman, J. R. Brister, and C. O’Sullivan. Stat: a fast, scalable, minhash- based k-mer tool to assess sequence read archive next-generation sequence submissions. Genome Biology, 22(1):1–15, 2021

  35. [43]

    S. M. Kiełbasa, R. Wan, K. Sato, P. Horton, and M. C. Frith. Adaptive seeds tame genomic sequence comparison. Genome research, 21(3):487–493, 2011

  36. [44]

    Kodama, M

    Y . Kodama, M. Shumway, and R. Leinonen. The sequence read archive: explosive growth of sequencing data. Nucleic acids research, 40(D1):D54–D56, 2012

  37. [45]

    G. R. Krause, W. Shands, and T. J. Wheeler. Sensitive and error-tolerant annotation of protein-coding dna with bath. Bioinformatics Advances, page vbae088, 2024

  38. [46]

    Krogh, M

    A. Krogh, M. Brown, I. S. Mian, K. Sjölander, and D. Haussler. Hidden markov models in computational biology: Applications to protein modeling. Journal of molecular biology, 235(5):1501–1531, 1994

  39. [47]

    T. W. Lab. Bagel: A framework for benchmarking sequence search tools. https://github.com/ TravisWheelerLab/BAGEL, 2024. Accessed: 2024-07-19

  40. [48]

    A. A. Larkin, C. A. Garcia, N. Garcia, M. L. Brock, J. A. Lee, L. J. Ustick, L. Barbero, B. R. Carter, R. E. Sonnerup, L. D. Talley, et al. High spatial resolution global ocean metagenomes from bio-go-ship repeat hydrography transects. Scientific data, 8(1):107, 2021

  41. [49]

    H. Li. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 34(18):3094–3100, 2018

  42. [50]

    J. Li, H. Jia, X. Cai, H. Zhong, Q. Feng, S. Sunagawa, M. Arumugam, J. R. Kultima, E. Prifti, T. Nielsen, et al. An integrated catalog of reference genes in the human gut microbiome. Nature biotechnology, 32(8):834–841, 2014

  43. [51]

    Liu and D

    S. Liu and D. Koslicki. Cmash: fast, multi-resolution estimation of k-mer-based jaccard and containment indices. Bioinformatics, 38(Supplement_1):i28–i35, 2022

  44. [52]

    B. Ma, M. T. France, J. Crabtree, J. B. Holm, M. S. Humphrys, R. M. Brotman, and J. Ravel. A comprehensive non-redundant gene catalog reveals extensive within-community intraspecies diversity in the human vagina. Nature communications, 11(1):940, 2020

  45. [53]

    P. B. McGarvey, A. Nightingale, J. Luo, H. Huang, M. J. Martin, C. Wu, and U. Consortium. Uniprot genomic mapping for deciphering functional effects of missense variants. Human mutation, 40(6):694–705, 2019

  46. [54]

    Mehringer, E

    S. Mehringer, E. Seiler, F. Droop, M. Darvish, R. Rahn, M. Vingron, and K. Reinert. Hierarchical interleaved bloom filter: enabling ultrafast, approximate sequence queries. Genome Biology, 24(1):131, 2023

  47. [55]

    Meyer, A

    F. Meyer, A. Fritz, Z.-L. Deng, D. Koslicki, T. R. Lesker, A. Gurevich, G. Robertson, M. Alser, D. Antipov, F. Beghini, et al. Critical assessment of metagenome interpretation: the second round of challenges. Nature methods, 19(4):429–440, 2022. 20 A PREPRINT

  48. [56]

    Mölder, K

    F. Mölder, K. P. Jablonski, B. Letcher, M. B. Hall, C. H. Tomkins-Tinch, V . Sochat, J. Forster, S. Lee, S. O. Twardziok, A. Kanitz, et al. Sustainable data analysis with snakemake. F1000Research, 10, 2021

  49. [57]

    About the STRIDES initiative

    National Institutes of Health. About the STRIDES initiative. https://datascience.nih.gov/strides

  50. [58]

    SRA end user cloud access costs

    National Institutes of Health. SRA end user cloud access costs. https://www.ncbi.nlm.nih.gov/sra/docs/ sra-cloud-access-costs/

  51. [59]

    Amazon web services joins NIH’s STRIDES initiative to harness latest cloud technologies for biomedical researchers, 2018

    National Institutes of Health. Amazon web services joins NIH’s STRIDES initiative to harness latest cloud technologies for biomedical researchers, 2018. https://www.nih.gov/news-events/news-releases/ amazon-web-services-joins-nihs-strides-initiative-harness-latest-cloud-techno...

  52. [60]

    NIH makes strides to accelerate discoveries in the cloud, 2018

    National Institutes of Health. NIH makes strides to accelerate discoveries in the cloud, 2018. https://www.nih. gov/news-events/news-releases/nih-makes-strides-accelerate-discoveries-cloud

  53. [61]

    NMDC metadata for soil study

    National Microbiome Data Collaborative. NMDC metadata for soil study. https://data.microbiomedata. org/

  54. [62]

    NMDC workflows

    National Microbiome Data Collaborative. NMDC workflows. https://nmdc-workflow-documentation. readthedocs.io/en/latest/chapters/overview.html#nmdc

  55. [63]

    E. P. Nawrocki, D. L. Kolbe, and S. R. Eddy. Infernal 1.0: inference of rna alignments. Bioinformatics, 25(10):1335–1337, 2009

  56. [64]

    H. B. Nielsen, M. Almeida, A. S. Juncker, S. Rasmussen, J. Li, S. Sunagawa, D. R. Plichta, L. Gautier, A. G. Pedersen, E. Le Chatelier, et al. Identification and assembly of genomes and genetic elements in complex metagenomic samples without using reference genomes. Nature bio...

  57. [65]

    S. Nurk, S. Koren, A. Rhie, M. Rautiainen, A. V . Bzikadze, A. Mikheenko, M. R. V ollger, N. Altemose, L. Uralsky, A. Gershman, et al. The complete sequence of a human genome. Science, 376(6588):44–53, 2022

  58. [66]

    B. D. Ondov, G. J. Starrett, A. Sappington, A. Kostic, S. Koren, C. B. Buck, and A. M. Phillippy. Mash screen: high-throughput sequence containment estimation for genome discovery. Genome biology, 20:1–13, 2019

  59. [67]

    B. D. Ondov, T. J. Treangen, P. Melsted, A. B. Mallonee, N. H. Bergman, S. Koren, and A. M. Phillippy. Mash: fast genome and metagenome distance estimation using minhash. Genome biology, 17:1–14, 2016

  60. [68]

    Pandey, F

    P. Pandey, F. Almodaresi, M. A. Bender, M. Ferdman, R. Johnson, and R. Patro. Mantis: a fast, small, and exact large-scale sequence-search index. Cell systems, 7(2):201–207, 2018

  61. [69]

    G. W. Park, T. F. F. Ng, A. L. Freeland, V . C. Marconi, J. A. Boom, M. A. Staat, A. M. Montmayeur, H. Browne, J. Narayanan, D. C. Payne, et al. Crassphage as a novel tool to detect human fecal contamination on environmental surfaces and hands. Emerging Infectious Diseases, 26...

  62. [70]

    D. H. Parks, C. Rinke, M. Chuvochina, P.-A. Chaumeil, B. J. Woodcroft, P. N. Evans, P. Hugenholtz, and G. W. Tyson. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nature microbiology, 2(11):1533–1542, 2017

  63. [71]

    N. T. Pierce, L. Irber, T. Reiter, P. Brooks, and C. T. Brown. Large-scale sequence comparisons with sourmash. F1000Research, 8, 2019

  64. [72]

    Pruitt, K

    K. Pruitt, K. Clark, T. Tatusova, and I. Mizrachi. Bioproject help. In BioProject Help [Internet]. National Center for Biotechnology Information (US), 2011

  65. [73]

    J. Qin, R. Li, J. Raes, M. Arumugam, K. S. Burgdorf, C. Manichanh, T. Nielsen, N. Pons, F. Levenez, T. Yamada, et al. A human gut microbial gene catalogue established by metagenomic sequencing. nature, 464(7285):59–65, 2010

  66. [74]

    J. W. Roddy, D. H. Rich, and T. J. Wheeler. nail: software for high-speed, high-sensitivity protein sequence annotation. bioRxiv, 2024

  67. [75]

    Seiler, S

    E. Seiler, S. Mehringer, M. Darvish, E. Turc, and K. Reinert. Raptor: A fast and space-efficient pre-filter for querying very large collections of nucleotide sequences. Iscience, 24(7), 2021

  68. [76]

    Sharon, M

    I. Sharon, M. J. Morowitz, B. C. Thomas, E. K. Costello, D. A. Relman, and J. F. Banfield. Time series community genomics analysis reveals rapid shifts in bacterial species, strains, and phage during infant gut colonization. Genome research, 23(1):111–120, 2013

  69. [77]

    S. A. Shiryev and R. Agarwala. Indexing and searching petabase-scale nucleotide resources. Nature methods, 21(6):994–1002, 2024

  70. [78]

    Shumway, G

    M. Shumway, G. Cochrane, and H. Sugawara. Archiving next generation sequencing data. Nucleic acids research, 38(suppl_1):D870–D871, 2010. 21 A PREPRINT

  71. [79]

    J. Söding. Protein homology detection by hmm–hmm comparison. Bioinformatics, 21(7):951–960, 2005

  72. [80]

    Solomon and C

    B. Solomon and C. Kingsford. Fast search of thousands of short-read sequencing experiments. Nature biotechnol- ogy, 34(3):300–302, 2016

  73. [81]

    branchwater: Searching large collections of sequencing data with genome-scale queries

    sourmash. branchwater: Searching large collections of sequencing data with genome-scale queries. https: //github.com/sourmash-bio/branchwater

  74. [82]

    Steinegger and J

    M. Steinegger and J. Söding. Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature biotechnology, 35(11):1026–1028, 2017

  75. [83]

    B. J. Tully, E. D. Graham, and J. F. Heidelberg. The reconstruction of 2,631 draft metagenome-assembled genomes from the global oceans. Scientific data, 5(1):1–8, 2018

  76. [84]

    P. J. Turnbaugh, R. E. Ley, M. Hamady, C. M. Fraser-Liggett, R. Knight, and J. I. Gordon. The human microbiome project. Nature, 449(7164):804–810, 2007

  77. [85]

    J. C. Venter, M. D. Adams, E. W. Myers, P. W. Li, R. J. Mural, G. G. Sutton, H. O. Smith, M. Yandell, C. A. Evans, R. A. Holt, et al. The sequence of the human genome. science, 291(5507):1304–1351, 2001

  78. [86]

    T. J. Wheeler, J. Clements, S. R. Eddy, R. Hubley, T. A. Jones, J. Jurka, A. F. Smit, and R. D. Finn. Dfam: a database of repetitive dna based on profile hidden markov models. Nucleic acids research, 41(D1):D70–D82, 2012

  79. [87]

    T. J. Wheeler and S. R. Eddy. nhmmer: Dna homology search with profile hmms. Bioinformatics, 29(19):2487– 2489, 2013

  80. [88]

    M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, et al. The FAIR guiding principles for scientific data management and stewardship. Scientific data, 3(1):1–9, 2016

  81. [89]

    D. E. Wood, J. Lu, and B. Langmead. Improved metagenomic analysis with kraken 2. Genome biology, 20:1–13, 2019

  82. [90]

    D. E. Wood and S. L. Salzberg. Kraken: ultrafast metagenomic sequence classification using exact alignments. Genome biology, 15:1–12, 2014

  83. [91]

    E. S. Wright. Using decipher v2. 0 to analyze big biological sequence data in r. R Journal, 8(1), 2016

  84. [92]

    Wu and Y

    S. Wu and Y . Zhang. Muster: improving protein sequence profile–profile alignments by using multiple sources of structure information. Proteins: Structure, Function, and Bioinformatics, 72(2):547–556, 2008

  85. [93]

    J. Ye, S. McGinnis, and T. L. Madden. Blast: improvements for better sequence analysis. Nucleic acids research, 34(suppl_2):W6–W9, 2006

  86. [94]

    Y . Yu, J. Liu, X. Liu, Y . Zhang, E. Magner, E. Lehnert, C. Qian, and J. Liu. Seqothello: querying rna-seq experiments at scale. Genome biology, 19:1–13, 2018. 7 Appendix 7.1 Appendix 1: Marine metagenome data accessions Ocean metagenome dataset used to generate gold standard...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.