Pith. sign in

REVIEW 1 major objections 1 minor 88 references

ChloroScan recovers plastid genomes from metagenomes by combining deep-learning contig filtering with marker-gene-guided binning, outperforming a generic binner on simulations and recovering 16 plastid genomes from four ocean metagenomes.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 10:12 UTC pith:CLNX5Y7C

load-bearing objection Useful pipeline with a solid simulation benchmark, but the real-data quality numbers hinge on the same marker database used for binning and QC, so they likely overstate completeness and purity. the 1 major comments →

arxiv 2510.10950 v1 pith:CLNX5Y7C submitted 2025-10-13 q-bio.GN

ChloroScan: Recovering plastid genome bins from metagenomic data

classification q-bio.GN
keywords plastid genomesmetagenome-assembled genomesmetagenomic binningdeep learning contig classificationmarker gene databaseprotist genomicsmarine metagenomesplastid MAGs
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

ChloroScan is an automated pipeline for pulling plastid genomes out of metagenomes, where nuclear protist genomes are often too complex to assemble but organellar genomes are small, high-copy, and phylogenetically informative. The paper argues that a deep-learning contig classifier followed by binning guided by a curated plastid marker-gene database recovers plastid metagenome-assembled genomes more completely and purely than a generic binner on simulated data. Applied to four ocean metagenomes, ChloroScan recovered 16 plastid metagenome-assembled genomes at 70 percent completeness and 90 percent purity, including a likely deep-branching marine ochrophyte lineage with no close sequenced relatives. If correct, this makes plastid genomes accessible from existing metagenomic libraries at scale, easing a bottleneck in discovering protist diversity.

Core claim

The paper's central claim is that plastid genome recovery from metagenomes improves when binning is guided by a plastid-specific marker-gene database rather than by generic prokaryotic markers. ChloroScan's pipeline first uses a deep-learning contig classifier to enrich for plastid contigs, then clusters those contigs with a marker-gene-guided binner, then adds taxonomic assignment and gene prediction. On two simulated marine metagenomes, it recovered more high-quality plastid bins than the benchmark binner, including near-complete single-contig genomes; its F1 and base-level accuracy were higher in both samples. On four real ocean metagenomes it produced 16 plastid metagenome-assembled geno

What carries the argument

The load-bearing mechanism is a manually curated database of plastid-encoded marker genes, formatted in the same style as prokaryotic quality-assessment marker sets, which steers the binning algorithm to cluster contigs into complete, pure plastid bins. Around this core, a deep-learning contig classifier pre-filters assemblies to a plastid-enriched set, and contig-level taxonomy is assigned by comparing predicted proteins against a combined reference protein database; a final round of homology searches on the rbcL marker provides fine-grained identification. The marker database and the classifier thresholds are the two settings that most directly determine which plastid lineages are recovere

Load-bearing premise

The pipeline assumes the deep-learning classifier's training set and the curated plastid marker-gene database cover the target lineage; unusual plastid architectures absent from those resources, such as dinoflagellate minicircles, are filtered out before binning and cannot be recovered.

What would settle it

Simulate a metagenome containing only high-coverage plastid genomes from lineages absent from the classifier's training data, for example dinoflagellate minicircle chromosomes, run ChloroScan, and check whether plastid bins emerge; if the genomes are in the assembly but no bin is recovered, the pipeline's coverage assumption is breached. A second check is to re-run the benchmark on the same simulated communities with several random seeds and see whether the reported F1 advantage over the generic binner persists.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Existing metagenomic libraries can be re-mined automatically for plastid genomes, without manual or human-guided binning.
  • Recovered plastid genomes carry phylogenetic markers such as rbcL and coding sequences, so they can feed directly into algal phylogenomics and species-delimitation studies.
  • Adapting the marker-gene database to mitochondrial genes would extend the same workflow to heterotrophic protists, which lack plastids.
  • As reference plastid genomes grow, both the classifier and marker database will cover more lineages, improving recovery and taxonomic resolution.
  • Low-coverage plastid genomes (roughly below 5x average depth) are likely to be missed, so the method's gains concentrate on abundant or deeply sequenced taxa.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit is that the success of marker-gene-guided binning for plastids could generalize to other organellar or extrachromosomal elements, such as mitochondria or plasmids, wherever conserved single-copy genes exist.
  • If ChloroScan scales to the full set of public metagenomes, it could substantially expand the sampled plastid tree of life; the 85 percent rbcL hit in this paper hints at how much novelty remains in under-sequenced marine lineages.
  • The systematic failure on dinoflagellate minicircle plastids implies that lineage-specific retraining of the classifier, not just more reference genomes, is needed; users should expect recovery to be biased toward lineages already represented in the training data.
  • A practical test would be applying ChloroScan to mock communities with known plastid abundances to quantify how the roughly 5x coverage threshold shifts with contig length cutoffs and binning parameters.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. ChloroScan is a Snakemake pipeline for recovering plastid genome bins from metagenomic assemblies. It first filters contigs with the deep-learning classifier Corgi, then performs marker-guided binning by adapting binny with a custom plastid marker-gene database, assigns taxonomy with CAT/BAT, and provides QC, annotation, and visualization outputs. The authors benchmark ChloroScan against MetaBAT2 on two CAMISIM-simulated marine metagenomes with known ground truth, reporting higher F1 scores, higher base-level accuracy, and more high-quality plastid bins, including several near-complete single-contig MAGs. They also apply ChloroScan to four Tara Oceans metagenomes and report 16 ptMAGs with completeness >70% and purity >90%, including a bin suggested to represent a novel ochrophyte lineage. The manuscript explicitly acknowledges that dinoflagellate minicircle plastid genomes were not recovered because Corgi's training set lacks such sequences.

Significance. If the results hold, ChloroScan would be a valuable, practical addition to the small set of tools targeting plastid genomes from metagenomes. The simulated benchmark is externally grounded: CAMISIM-generated reads with known genomes and AMBER assessment avoid circularity in the MetaBAT2 comparison. The code, reproducibility scripts, and intermediate data are made available, which is a clear strength. However, the real-data quality claims are weakened by the circular use of the same custom marker database both to guide binning and to compute marker-based completeness/purity, and the headline count of 16 ptMAGs appears to include a bin later identified as a likely contaminant. These are fixable with additional validation and recalculation, so the manuscript is a candidate for major revision rather than rejection.

major comments (1)
  1. [§2.1, §3.2] The real-data quality metrics are circular. The custom plastid marker database is used by binny to guide clustering, and the same database is then used in the summary/QC module to count marker genes, compute marker completeness, and flag contaminant contigs described as 'without ORFs matching our plastid marker gene database'. Bins assembled around marker-bearing contigs can therefore pass marker-based completeness thresholds by construction, while genuine plastid contigs from divergent lineages lacking these markers may be removed, inflating purity. This does not affect the simulated benchmark, which uses AMBER ground truth, but it directly undermines the claimed >70% completeness and >90% purity for the 16 real ptMAGs and their MIMAG classification. Please validate the real bins with an independent approach—for example, split-marker or leave-one-out cross-validation, mapping/coverage e
minor comments (1)
  1. [Figure 4] The caption states 'six bins from the sample SAMEA2732360', but the text describes eight bins (Bin 0 through Bin 7) for that sample. Please harmonize the caption and text.

Circularity Check

1 steps flagged

Real-data ptMAG quality partly reduces to the same marker database used to steer binning; the synthetic benchmark remains externally grounded.

specific steps
  1. self definitional [§2.1 (binning and QC design); §3.2 (real-metagenome result)]
    "we designed a custom database for plastid marker genes ... transforming it to recover plastid bins. ... we adopted a conservative contamination detection ... without ORFs matching our plastid marker gene database. ... the binning module inferred 16 ptMAGs with completeness > 70% and purity > 90%"

    The custom plastid marker-gene database is used both to guide binny's clustering of contigs into plastid bins and to compute marker-based completeness and to flag contaminant contigs lacking those markers. A bin is thus assembled preferentially around contigs that carry the very markers later counted for quality assessment. Consequently, the real-metagenome claim of 16 ptMAGs with >70% completeness and >90% purity is partly a restatement of the binning objective rather than an independent quality measurement. The simulated benchmark is not affected, because AMBER uses ground-truth mapping to source genomes.

full rationale

The central benchmark comparison against MetaBAT2 on simulated metagenomes is externally grounded: ChloroScan bins are evaluated with AMBER against known source genomes, so the claim that ChloroScan recovers more high-quality plastid bins on synthetic data does not reduce to the tool's own quality scoring. The main circularity lies in the real-data analysis: the same manually curated plastid marker database is used by binny to guide clustering and by the summary/QC module to count marker genes, estimate marker completeness, and filter contaminant contigs; therefore the '16 ptMAGs with completeness > 70% and purity > 90%' result is in part constructed from the same marker set used to define the bins. The paper's acknowledged dinoflagellate minicircle failure reflects a training-data limitation of Corgi rather than a circular step, and the Corgi self-citation is not load-bearing because the pipeline is independently benchmarked on simulated data. Overall, the independent synthetic benchmark keeps the score moderate, but the real-data quality assertion is partially circular and should be supported by split-marker or external validation.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 0 invented entities

No new biological entities are postulated. The custom plastid marker database is a constructed resource, not an invented entity. The central claim rests on several hand-set thresholds and domain assumptions about the transferability of prokaryote-oriented binning and quality-control tools to plastid genomes.

free parameters (6)
  • Corgi plastid probability threshold = 0.50 (default)
    Contigs assigned to plastid category if plastid probability exceeds all others and is >=0.50 (Section 2.1). Hand-set, no sensitivity analysis.
  • Contig length cutoff = 1000 bp default; 1500 bp in benchmark
    Short contigs filtered before binning; cutoff affects recovery of fragmented plastids (Sections 2.1, 2.2).
  • Marker gene count threshold for scMAGs = 30
    Single-contig MAGs require at least 30 marker genes (Section 2.1). Hand-set.
  • Marker completeness threshold for scMAGs = 85%
    Single-contig MAGs require 85% marker completeness (Section 2.1). Hand-set.
  • Bin completeness and purity cutoffs for real data = completeness 70%, purity 90%
    Used to select medium-to-high-quality bins in real metagenomes (Section 2.3). Hand-set and not benchmarked across datasets.
  • Minimum completeness in simulations = 50%
    Set to recover as many bins as possible in simulated data (Section 2.2).
axioms (5)
  • domain assumption Corgi's classifier probability for 'plastid' category correctly reflects biological origin across diverse plastid genomes.
    The whole pipeline depends on this classifier to enrich for plastid contigs; the paper explicitly shows failure for dinoflagellate minicircles (Section 2.1, Discussion).
  • domain assumption Binny's marker-gene-guided clustering algorithm remains effective when the marker database is replaced with plastid-specific markers.
    Binny was designed for prokaryotes; using a custom plastid marker set assumes the algorithm generalizes (Section 2.1).
  • domain assumption MIMAG completeness and purity thresholds, developed for microbial genomes, are meaningful quality measures for plastid MAGs.
    The paper applies MIMAG categories to plastid bins (Sections 2.3, 3.1) without validating that these thresholds behave equivalently for plastid genomes.
  • domain assumption CAMISIM-simulated metagenomes capture the assembly and binning complexity of real marine metagenomes.
    The central benchmark conclusion is based on two simulated samples; generalization to real data relies on this assumption (Section 2.2).
  • domain assumption CAT/BAT taxonomic predictions with the custom protein database are sufficiently accurate for contamination detection and taxonomic placement.
    Contamination detection and taxonomic interpretation of bins depend on CAT/BAT results (Sections 2.1, 3.2).

pith-pipeline@v1.3.0-alltime-deepseek · 16081 in / 8955 out tokens · 78707 ms · 2026-08-04T10:12:04.029531+00:00 · methodology

0 comments
read the original abstract

Genome-resolved metagenomics has contributed largely to discovering prokaryotic genomes. When applied to microscopic eukaryotes, challenges such as the high number of introns and repeat regions found in nuclear genomes have hampered the mining and discovery of novel protistan lineages. Organellar genomes are simpler, smaller, have higher abundance than their nuclear counterparts and contain valuable phylogenetic information, but are yet to be widely used to identify new protist lineages from metagenomes. Here we present "ChloroScan", a new bioinformatics pipeline to extract eukaryotic plastid genomes from metagenomes. It incorporates a deep learning contig classifier to identify putative plastid contigs and an automated binning module to recover bins with guidance from a curated marker gene database. Additionally, ChloroScan summarizes the results in different user-friendly formats, including annotated coding sequences and proteins for each bin. We show that ChloroScan recovers more high-quality plastid bins than MetaBAT2 for simulated metagenomes. The practical utility of ChloroScan is illustrated by recovering 16 medium to high-quality metagenome assembled genomes from four protist-size fractioned metagenomes, with several bins showing high taxonomic novelty.

Figures

Figures reproduced from arXiv: 2510.10950 by Heroen Verbruggen, Robert Turnbull, Vanessa Rossetto Marcelino, Yuhao Tong.

Figure 1
Figure 1. Figure 1: ChloroScan’s workflow structure. ChloroScan contains the following modules: contig prediction by Corgi [50], binning by binny [47], taxonomic prediction by CAT/BAT [51], and a summary module that generates user-friendly information to investigate the contig and bin data, including plots to investigate bin homogeneity, a table with contig metadata and the predicted genes and proteins from MAGs. It takes ass… view at source ↗
Figure 2
Figure 2. Figure 2: Plots of ChloroScan results show its effectiveness compared to MetaBAT2. The bar charts compare bins from ChloroScan to bins from MetaBAT2 in (a) single sample metagenome 1 and (b) single sample metagenome 2, in terms of their homogeneity by showing how many taxa are included in one bin. ChloroScan bins are labelled as digits and MetaBAT2 bins are prepended with “meta”. Single contig MAGs recovered by Chlo… view at source ↗
Figure 3
Figure 3. Figure 3: Mapping information from each bin to source genomes in the (a) synthetic single￾sample metagenome 1 and (b) 2 based on the contig mapping information generated from CAMISIM. Grey bar widths refer to the percentage of source genome length taken by contigs. Here one contig has only one source genome mapped. Source genomes with too short contigs (colors invisible in the Figure) in the sample have their names … view at source ↗
Figure 4
Figure 4. Figure 4: Metagenome-assembled genomes from real marine metagenomes. a. The GC x log10 average read depth plots of the sample SAMEA2732360. Marker gene count per contig is scaled by the dot size. b. Contig-level taxonomy composition of six bins from the sample SAMEA2732360 inferred by CAT. The sorted percentages of MAG length taken by each taxon are listed on the left side, and the detailed taxon lineages (adapted f… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

88 extracted references · 61 canonical work pages

  1. [1]

    Melbourne Integrative Genomics, School of BioSciences, University of Melbourne, Melbourne, VIC 3010, Australia

  2. [2]

    Institute of Agrochemistry and Food Technology, Spanish National Research Council (CSIC), Valencia, 46980, Spain

  3. [3]

    Department of Microbiology and Immunology at the Peter Doherty Institute for Infection and Immunity, University of Melbourne, Melbourne, VIC 3010, Australia

  4. [4]

    Melbourne Data Analytics Platform (MDAP), Melbourne Connect, University of Melbourne, Melbourne, VIC, Australia

  5. [5]

    ChloroScan

    CIBIO, Centro de Investigaç ã o em Biodiversidade e Recursos Gené ticos, InBIO Laborató rio Associado, Campus de Vairã o, Universidade do Porto, 4485-661 Vairã o, Portugal. Contact: yuhtong@student.unimelb.edu.au Abstract Genome-resolved metagenomics has contributed largely to discovering prokaryotic genomes. When applied to microscopic eukaryotes (protis...

  6. [6]

    Introduction The amount of sequenced data for microbiomes has ballooned in the last decade [1]. Genome-resolved metagenomics (GRM) became a widely used approach to analyze these data for environmental microbiomes, offering numerous insights of their evolution, ecology and diversity [2]. It incorporates de novo assembly to build contigs and binning algorit...

  7. [7]

    It is underpinned by a deep learning-based module to predict contigs of plastid origin and a manually curated database to guide metagenome binning and quality assessment (Figure 1)

    Materials and Methods 2.1 ChloroScan workflow overview ChloroScan is a Snakemake-based [48] workflow with the command line interface wrapped by snk [49], to infer ptMAGs from metagenome contigs, with the utilities mostly written in Python and Unix bash. It is underpinned by a deep learning-based module to predict contigs of plastid origin and a manually c...

  8. [8]

    L. A. Hug et al., ‘A new view of the tree of life’, Nat. Microbiol., vol. 1, no. 5, p. 16048, Apr. 2016, doi: 10.1038/nmicrobiol.2016.48

  9. [9]

    Results 3.1 ChloroScan recovers high-quality ptMAGs from synthetic metagenomes To compare ChloroScan’s performance against that of similar software, we benchmarked it alongside MetaBAT2, the binner used by plastiC [42]. Both solutions produced plastid MAGs (Figure 2), but we found that those produced with ChloroScan had higher overall quality and purity, ...

  10. [10]

    uncultured bacterium

    Discussion We developed ChloroScan, a metagenomic binning workflow targeting plastid genomes, and showed its performance using synthetic and real metagenomes. ChloroScan leverages an existing binning framework designed for prokaryotes [47], but we enhanced its performance for plastid genome binning with a manually designed plastid- encoded marker gene dat...

  11. [11]

    Quince, A

    C. Quince, A. W. Walker, J. T. Simpson, N. J. Loman, and N. Segata, ‘Shotgun metagenomics, from sampling to analysis’, Nat. Biotechnol., vol. 35, no. 9, pp. 833–844, Sept. 2017, doi: 10.1038/nbt.3935

  12. [12]

    V. W. Salazar et al., ‘Metaphor—A workflow for streamlined assembly and binning of metagenomes’, GigaScience, vol. 12, p. giad055, Dec. 2022, doi: 10.1093/gigascience/giad055

  13. [13]

    Nayfach et al., ‘A genomic catalog of Earth’s microbiomes’, Nat

    S. Nayfach et al., ‘A genomic catalog of Earth’s microbiomes’, Nat. Biotechnol., vol. 39, no. 4, pp. 499–509, Apr. 2021, doi: 10.1038/s41587-020-0718-6

  14. [14]

    Xiao et al., ‘Microbial ecosystems and ecological driving forces in the deepest ocean sediments’, Cell, vol

    X. Xiao et al., ‘Microbial ecosystems and ecological driving forces in the deepest ocean sediments’, Cell, vol. 188, no. 5, pp. 1363-1377.e9, Mar. 2025, doi: 10.1016/j.cell.2024.12.036

  15. [15]

    Dai et al., ‘Crop root bacterial and viral genomes reveal unexplored species and microbiome patterns’, Cell, vol

    R. Dai et al., ‘Crop root bacterial and viral genomes reveal unexplored species and microbiome patterns’, Cell, vol. 188, no. 9, pp. 2521-2539.e22, May 2025, doi: 10.1016/j.cell.2025.02.013

  16. [16]

    Yang et al., ‘Metagenomic and metatranscriptomic analyses reveal minor-yet- crucial roles of gut microbiome in deep-sea hydrothermal vent snail’, Anim

    Y. Yang et al., ‘Metagenomic and metatranscriptomic analyses reveal minor-yet- crucial roles of gut microbiome in deep-sea hydrothermal vent snail’, Anim. Microbiome, vol. 4, no. 1, p. 3, Dec. 2022, doi: 10.1186/s42523-021-00150-z

  17. [17]

    C. T. Brown et al., ‘Unusual biology across a group comprising more than 15% of domain Bacteria’, Nature, vol. 523, no. 7559, pp. 208–211, July 2015, doi: 10.1038/nature14486

  18. [18]

    Carradec et al., ‘A global ocean atlas of eukaryotic genes’, Nat

    Q. Carradec et al., ‘A global ocean atlas of eukaryotic genes’, Nat. Commun., vol. 9, no. 1, p. 373, Jan. 2018, doi: 10.1038/s41467-017-02342-1

  19. [19]

    Liu et al., ‘Expanded diversity of Asgard archaea and their relationships with eukaryotes’, Nature, vol

    Y. Liu et al., ‘Expanded diversity of Asgard archaea and their relationships with eukaryotes’, Nature, vol. 593, no. 7860, pp. 553–557, May 2021, doi: 10.1038/s41586- 021-03494-3

  20. [20]

    Leã o et al., ‘Asgard archaea defense systems and their roles in the origin of eukaryotic immunity’, Nat

    P. Leã o et al., ‘Asgard archaea defense systems and their roles in the origin of eukaryotic immunity’, Nat. Commun., vol. 15, no. 1, p. 6386, July 2024, doi: 10.1038/s41467-024-50195-2

  21. [21]

    We altered binny to not run read depth calculations, but rather do this more efficiently within the core ChloroScan workflow

    (Supplementary Materials), transforming it to recover plastid bins. We altered binny to not run read depth calculations, but rather do this more efficiently within the core ChloroScan workflow. For the recovery of single-contig MAGs (scMAGs), we set default thresholds of 30 for marker gene count and 85% for marker completeness. Following binning, taxonomi...

  22. [22]

    Eme and D

    L. Eme and D. Tamarit, ‘Microbial Diversity and Open Questions about the Deep Tree of Life’, Genome Biol. Evol., vol. 16, no. 4, p. evae053, Apr. 2024, doi: 10.1093/gbe/evae053

  23. [23]

    J. M. Diaz and S. Plummer, ‘Production of extracellular reactive oxygen species by phytoplankton: past and future directions’, J. Plankton Res., Sept. 2018, doi: 10.1093/plankt/fby039

  24. [24]

    Solomon et al., ‘Protozoa populations are ecosystem engineers that shape prokaryotic community structure and function of the rumen microbial ecosystem’, ISME J., vol

    R. Solomon et al., ‘Protozoa populations are ecosystem engineers that shape prokaryotic community structure and function of the rumen microbial ecosystem’, ISME J., vol. 16, no. 4, pp. 1187–1197, Apr. 2022, doi: 10.1038/s41396-021-01170-y

  25. [25]

    Chabé, A

    M. Chabé, A. Lokmer, and L. Ségurel, ‘Gut Protozoa: Friends or Foes of the Human Gut Microbiota?’, Trends Parasitol., vol. 33, no. 12, pp. 925–934, Dec. 2017, doi: 10.1016/j.pt.2017.08.005

  26. [26]

    S. J. Sibbald and J. M. Archibald, ‘More protist genomes needed’, Nat. Ecol. Evol., vol. 1, no. 5, p. 0145, Apr. 2017, doi: 10.1038/s41559-017-0145

  27. [27]

    Miao et al., ‘Protist 10,000 Genomes Project’, The Innovation, vol

    W. Miao et al., ‘Protist 10,000 Genomes Project’, The Innovation, vol. 1, no. 3, p. 100058, Nov. 2020, doi: 10.1016/j.xinn.2020.100058

  28. [28]

    Gao et al., ‘The P10K database: a data portal for the protist 10 000 genomes project’, Nucleic Acids Res., vol

    X. Gao et al., ‘The P10K database: a data portal for the protist 10 000 genomes project’, Nucleic Acids Res., vol. 52, no. D1, pp. D747–D755, Jan. 2024, doi: 10.1093/nar/gkad992

  29. [29]

    Alexander et al., ‘Eukaryotic genomes from a global metagenomic data set illuminate trophic modes and biogeography of ocean plankton’, mBio, vol

    H. Alexander et al., ‘Eukaryotic genomes from a global metagenomic data set illuminate trophic modes and biogeography of ocean plankton’, mBio, vol. 14, no. 6, pp. e01676-23, Dec. 2023, doi: 10.1128/mbio.01676-23

  30. [30]

    Laforest-Lapointe and M.-C

    I. Laforest-Lapointe and M.-C. Arrieta, ‘Microbial Eukaryotes: a Missing Link in Gut Microbiome Studies’, mSystems, vol. 3, no. 2, pp. e00201-17, Apr. 2018, doi: 10.1128/mSystems.00201-17

  31. [31]

    D. H. Parks, M. Imelfort, C. T. Skennerton, P. Hugenholtz, and G. W. Tyson, ‘CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes’, Genome Res., vol. 25, no. 7, pp. 1043–1055, July 2015, doi: 10.1101/gr.186072.114

  32. [32]

    Chklovski, D

    A. Chklovski, D. H. Parks, B. J. Woodcroft, and G. W. Tyson, ‘CheckM2: a rapid, scalable and accurate tool for assessing microbial genome quality using machine learning’, Nat. Methods, vol. 20, no. 8, pp. 1203–1212, Aug. 2023, doi: 10.1038/s41592- 023-01940-w

  33. [33]

    Chaumeil, A

    P.-A. Chaumeil, A. J. Mussig, P. Hugenholtz, and D. H. Parks, ‘GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database’, Bioinformatics, vol. 36, no. 6, pp. 1925–1927, Mar. 2020, doi: 10.1093/bioinformatics/btz848

  34. [34]

    Pan, X.-M

    S. Pan, X.-M. Zhao, and L. P. Coelho, ‘SemiBin2: self-supervised contrastive learning leads to better MAGs for short- and long-read sequencing’, Bioinformatics, vol. 39, no. Supplement_1, pp. i21–i29, June 2023, doi: 10.1093/bioinformatics/btad209

  35. [35]

    Z. Wang, R. You, H. Han, W. Liu, F. Sun, and S. Zhu, ‘Effective binning of metagenomic contigs using contrastive multi-view representation learning’, Nat. Commun., vol. 15, no. 1, p. 585, Jan. 2024, doi: 10.1038/s41467-023-44290-z

  36. [36]

    Kutuzova, M

    S. Kutuzova, M. Nielsen, P. Piera, J. N. Nissen, and S. Rasmussen, ‘Taxometer: Improving taxonomic classification of metagenomics contigs’, Nat. Commun., vol. 15, no. 1, p. 8357, Sept. 2024, doi: 10.1038/s41467-024-52771-y

  37. [37]

    T. O. Delmont et al., ‘Functional repertoire convergence of distantly related eukaryotic plankton lineages abundant in the sunlit ocean’, Cell Genomics, vol. 2, no. 5, p. 100123, May 2022, doi: 10.1016/j.xgen.2022.100123

  38. [38]

    N. V. Patin and K. D. Goodwin, ‘Long-Read Sequencing Improves Recovery of Picoeukaryotic Genomes and Zooplankton Marker Genes from Marine Metagenomes’, mSystems, vol. 7, no. 6, pp. e00595-22, Dec. 2022, doi: 10.1128/msystems.00595-22

  39. [39]

    Duncan et al., ‘Metagenome-assembled genomes of phytoplankton microbiomes from the Arctic and Atlantic Oceans’, Microbiome, vol

    A. Duncan et al., ‘Metagenome-assembled genomes of phytoplankton microbiomes from the Arctic and Atlantic Oceans’, Microbiome, vol. 10, no. 1, p. 67, Apr. 2022, doi: 10.1186/s40168-022-01254-7

  40. [40]

    A. I. Krinos, R. M. Bowers, R. R. Rohwer, K. D. McMahon, T. Woyke, and F. Schulz, ‘Time-series metagenomics reveals changing protistan ecology of a temperate dimictic lake’, Microbiome, vol. 12, no. 1, p. 133, July 2024, doi: 10.1186/s40168-024- 01831-y

  41. [41]

    S. J. Sibbald and J. M. Archibald, ‘Genomic Insights into Plastid Evolution’, Genome Biol. Evol., vol. 12, no. 7, pp. 978–990, July 2020, doi: 10.1093/gbe/evaa096

  42. [42]

    Piganeau and H

    G. Piganeau and H. Moreau, ‘Screening the Sargasso Sea metagenome for data to investigate genome evolution in Ostreococcus (Prasinophyceae, Chlorophyta)’, Gene, vol. 406, no. 1–2, pp. 184–190, Dec. 2007, doi: 10.1016/j.gene.2007.09.015

  43. [43]

    S. D. Gallaher, S. T. Fitz‐Gibbon, D. Strenkert, S. O. Purvine, M. Pellegrini, and S. S. Merchant, ‘High‐throughput sequencing of the chloroplast and mitochondrion of Chlamydomonas reinhardtii to generate improved de novo assemblies, analyze expression patterns and transcript speciation, and evaluate diversity among laboratory strains and wild isolates’, ...

  44. [44]

    Turmel, C

    M. Turmel, C. Otis, and C. Lemieux, ‘Dynamic Evolution of the Chloroplast Genome in the Green Algal Classes Pedinophyceae and Trebouxiophyceae’, Genome Biol. Evol., vol. 7, no. 7, pp. 2062–2082, July 2015, doi: 10.1093/gbe/evv130

  45. [45]

    Sauvage, W

    T. Sauvage, W. E. Schmidt, S. Suda, and S. Fredericq, ‘A metabarcoding framework for facilitated survey of endolithic phototrophs with tufA’, BMC Ecol., vol. 16, no. 1, p. 8, Dec. 2016, doi: 10.1186/s12898-016-0068-x

  46. [46]

    Borer, C

    G. Borer, C. Monteiro, F. P. Lima, and F. M. S. Martins, ‘Performance of DNA Metabarcoding vs. Morphological Methods for Assessing Intertidal Turf and Foliose Algae Diversity’, Mol. Ecol. Resour., vol. 25, no. 7, p. e14115, Oct. 2025, doi: 10.1111/1755-0998.14115

  47. [47]

    L. Sun, L. Fang, Z. Zhang, X. Chang, D. Penny, and B. Zhong, ‘Chloroplast Phylogenomic Inference of Green Algae Relationships’, Sci. Rep., vol. 6, no. 1, p. 20528, Feb. 2016, doi: 10.1038/srep20528

  48. [48]

    J. F. Costa, S.-M. Lin, E. C. Macaya, C. Ferná ndez-Garcí a, and H. Verbruggen, ‘Chloroplast genomes as a tool to resolve red algal phylogenies: a case study in the Nemaliales’, BMC Evol. Biol., vol. 16, no. 1, p. 205, Dec. 2016, doi: 10.1186/s12862- 016-0772-3

  49. [49]

    L. Fang, F. Leliaert, Z. Zhang, D. Penny, and B. Zhong, ‘Evolution of the Chlorophyta: Insights from chloroplast phylogenomic analyses’, J. Syst. Evol., vol. 55, no. 4, pp. 322–332, July 2017, doi: 10.1111/jse.12248

  50. [50]

    Accessed: Oct

    ‘National Center for Biotechnology Information’. Accessed: Oct. 08, 2025. [Online]. Available: https://www.ncbi.nlm.nih.gov/

  51. [51]

    M. D. Guiry, ‘How many species of algae are there? A reprise. Four kingdoms, 14 phyla, 63 classes and still growing’, J. Phycol., vol. 60, no. 2, pp. 214–228, Apr. 2024, doi: 10.1111/jpy.13431

  52. [52]

    E. S. Cameron, M. L. Blaxter, and R. D. Finn, ‘plastiC: A pipeline for recovery and characterization of plastid genomes from metagenomic datasets’, Wellcome Open Res., vol. 8, p. 475, Oct. 2023, doi: 10.12688/wellcomeopenres.19589.1

  53. [53]

    D. D. Kang et al., ‘MetaBAT 2: an adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies’, PeerJ, vol. 7, p. e7359, July 2019, doi: 10.7717/peerj.7359

  54. [54]

    Karlicki, S

    M. Karlicki, S. Antonowicz, and A. Karnkowska, ‘Tiara: deep learning-based classification system for eukaryotic sequences’, Bioinformatics, vol. 38, no. 2, pp. 344– 350, Jan. 2022, doi: 10.1093/bioinformatics/btab672

  55. [55]

    A. M. Eren et al., ‘Anvi’o: an advanced analysis and visualization platform for ‘omics data’, PeerJ, vol. 3, p. e1319, Oct. 2015, doi: 10.7717/peerj.1319

  56. [56]

    Karlicki, ‘Diversity and ecology of photosynthetic microbial eukaryotes in selected aquatic systems based on metabarcoding and metagenomic data’, Oct

    M. Karlicki, ‘Diversity and ecology of photosynthetic microbial eukaryotes in selected aquatic systems based on metabarcoding and metagenomic data’, Oct. 2024, Accessed: Oct. 08, 2025. [Online]. Available: https://repozytorium.uw.edu.pl//handle/item/160570

  57. [57]

    Hickl, P

    O. Hickl, P. Queiró s, P. Wilmes, P. May, and A. Heintz-Buschart, ‘binny : an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets’, Brief. Bioinform., vol. 23, no. 6, p. bbac431, Nov. 2022, doi: 10.1093/bib/bbac431

  58. [58]

    Mö lder et al., ‘Sustainable data analysis with Snakemake’, F1000Research, vol

    F. Mö lder et al., ‘Sustainable data analysis with Snakemake’, F1000Research, vol. 10, p. 33, Jan. 2021, doi: 10.12688/f1000research.29032.1

  59. [59]

    Wirth, S

    W. Wirth, S. Mutch, and R. Turnbull, ‘Snk: A Snakemake CLI and Workflow Management System’, J. Open Source Softw., vol. 9, no. 103, p. 7410, Nov. 2024, doi: 10.21105/joss.07410

  60. [60]

    Turnbull, rbturnbull/corgi

    R. Turnbull, rbturnbull/corgi. (Sept. 16, 2025). Python. Accessed: Oct. 08, 2025. [Online]. Available: https://github.com/rbturnbull/corgi

  61. [61]

    F. A. B. Von Meijenfeldt, K. Arkhipova, D. D. Cambuy, F. H. Coutinho, and B. E. Dutilh, ‘Robust taxonomic classification of uncharted microbial sequences and bins with CAT and BAT’, Genome Biol., vol. 20, no. 1, p. 217, Dec. 2019, doi: 10.1186/s13059- 019-1817-x

  62. [62]

    B. E. Suzek, H. Huang, P. McGarvey, R. Mazumder, and C. H. Wu, ‘UniRef: comprehensive and non-redundant UniProt reference clusters’, Bioinformatics, vol. 23, no. 10, pp. 1282–1288, May 2007, doi: 10.1093/bioinformatics/btm098

  63. [63]

    B. D. Ondov, N. H. Bergman, and A. M. Phillippy, ‘Interactive metagenomic visualization in a Web browser’, BMC Bioinformatics, vol. 12, no. 1, p. 385, Dec. 2011, doi: 10.1186/1471-2105-12-385

  64. [64]

    Van Der Jeugt, P

    F. Van Der Jeugt, P. Dawyndt, and B. Mesuere, ‘FragGeneScanRs: faster gene prediction for short reads’, BMC Bioinformatics, vol. 23, no. 1, p. 198, Dec. 2022, doi: 10.1186/s12859-022-04736-5

  65. [65]

    Fritz et al., ‘CAMISIM: simulating metagenomes and microbial communities’, Microbiome, vol

    A. Fritz et al., ‘CAMISIM: simulating metagenomes and microbial communities’, Microbiome, vol. 7, no. 1, p. 17, Dec. 2019, doi: 10.1186/s40168-019-0633-6

  66. [66]

    Accessed: Oct

    ‘GitHub - CAMI-challenge/CAMISIM at dev’, GitHub. Accessed: Oct. 08, 2025. [Online]. Available: https://github.com/CAMI-challenge/CAMISIM

  67. [67]

    Li, ‘Minimap2: pairwise alignment for nucleotide sequences’, Bioinformatics, vol

    H. Li, ‘Minimap2: pairwise alignment for nucleotide sequences’, Bioinformatics, vol. 34, no. 18, pp. 3094–3100, Sept. 2018, doi: 10.1093/bioinformatics/bty191

  68. [68]

    Meyer et al., ‘AMBER: Assessment of Metagenome BinnERs’, GigaScience, vol

    F. Meyer et al., ‘AMBER: Assessment of Metagenome BinnERs’, GigaScience, vol. 7, no. 6, p. giy069, June 2018, doi: 10.1093/gigascience/giy069

  69. [69]

    S. T. N. Aroney, R. J. P. Newell, J. N. Nissen, A. P. Camargo, G. W. Tyson, and B. J. Woodcroft, ‘CoverM: read alignment statistics for metagenomics’, Bioinformatics, vol. 41, no. 4, p. btaf147, Mar. 2025, doi: 10.1093/bioinformatics/btaf147

  70. [70]

    T. S. B. Schmidt et al., ‘SPIRE: a Searchable, Planetary-scale mIcrobiome REsource’, Nucleic Acids Res., vol. 52, no. D1, pp. D777–D783, Jan. 2024, doi: 10.1093/nar/gkad943

  71. [71]

    Biotechnol., vol

    The Genome Standards Consortium et al., ‘Minimum information about a single amplified genome (MISAG) and a metagenome-assembled genome (MIMAG) of bacteria and archaea’, Nat. Biotechnol., vol. 35, no. 8, pp. 725–731, Aug. 2017, doi: 10.1038/nbt.3893

  72. [72]

    J. L. Steenwyk and A. Rokas, ‘orthofisher: a broadly applicable tool for automated gene identification and retrieval’, G3 GenesGenomesGenetics, vol. 11, no. 9, p. jkab250, Sept. 2021, doi: 10.1093/g3journal/jkab250

  73. [73]

    J. P. Warnock, R. P. Scherer, and M. A. Konfirst, ‘A record of Pleistocene diatom preservation from the Amundsen Sea, West Antarctica with possible implications on silica leakage’, Mar. Micropaleontol., vol. 117, pp. 40–46, May 2015, doi: 10.1016/j.marmicro.2015.04.001

  74. [74]

    Zhang, Z

    M. Zhang, Z. Cui, F. Liu, and N. Chen, ‘Complete chloroplast genome of Eucampia zodiacus (Mediophyceae, Bacillariophyta)’, Mitochondrial DNA Part B, vol. 6, no. 8, pp. 2194–2197, Aug. 2021, doi: 10.1080/23802359.2021.1944828

  75. [75]

    Jamy et al., ‘New deep-branching environmental plastid genomes on the algal tree of life’, Jan

    M. Jamy et al., ‘New deep-branching environmental plastid genomes on the algal tree of life’, Jan. 2025, doi: 10.1101/2025.01.16.633336

  76. [76]

    V. R. Marcelino et al., ‘CCMetagen: comprehensive and accurate identification of eukaryotes and prokaryotes in metagenomic data’, Genome Biol., vol. 21, no. 1, p. 103, Dec. 2020, doi: 10.1186/s13059-020-02014-2

  77. [77]

    De Vargas et al., ‘Eukaryotic plankton diversity in the sunlit ocean’, Science, vol

    C. De Vargas et al., ‘Eukaryotic plankton diversity in the sunlit ocean’, Science, vol. 348, no. 6237, p. 1261605, May 2015, doi: 10.1126/science.1261605

  78. [78]

    J. J. Pierella Karlusich et al., ‘A robust approach to estimate relative phytoplankton cell abundances from metagenomes’, Mol. Ecol. Resour., vol. 23, no. 1, pp. 16–40, Jan. 2023, doi: 10.1111/1755-0998.13592

  79. [79]

    Penot, J

    M. Penot, J. B. Dacks, B. Read, and R. G. Dorrell, ‘Genomic and meta-genomic insights into the functions, diversity and global distribution of haptophyte algae’, Appl. Phycol., vol. 3, no. 1, pp. 340–359, Dec. 2022, doi: 10.1080/26388081.2022.2103732

  80. [80]

    Lopes Dos Santos et al., ‘Diversity and oceanic distribution of prasinophytes clade VII, the dominant group of green algae in oceanic waters’, ISME J., vol

    A. Lopes Dos Santos et al., ‘Diversity and oceanic distribution of prasinophytes clade VII, the dominant group of green algae in oceanic waters’, ISME J., vol. 11, no. 2, pp. 512–528, Feb. 2017, doi: 10.1038/ismej.2016.120

Showing first 80 references.