Pith. sign in

REVIEW 3 major objections 6 minor 32 references

A Transformer originally trained on microbial biosynthetic gene clusters can, after label-free adaptation on unlabeled plant genomes and weak supervision from functional annotations, discover plant BGCs without any curated plant labels, rec

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 16:11 UTC pith:X2DIFFDZ

load-bearing objection First learning-based plant BGC discovery pipeline worth taking seriously, but the headline false-positive reduction is circular until tested on labels other than the ones that trained it. the 3 major comments →

arxiv 2607.27258 v1 pith:X2DIFFDZ submitted 2026-07-29 q-bio.GN cs.LG

PlantBGC: Transformer for Plant BGC Discovery via Label-Free Domain Adaptation and Weak Supervision

classification q-bio.GN cs.LG
keywords biosynthetic gene clustersplant specialized metabolismgenome miningtransformerdomain adaptationweak supervisionmasked language modelingfalse-positive control
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that the scarcity of curated plant biosynthetic gene cluster (BGC) labels can be sidestepped by transferring supervision from well-annotated microbial BGCs to plants. It argues that an encoder-only Transformer, trained on ordered conserved-domain sequences from microbial BGCs and then adapted to unlabeled plant genomes via masked language modeling, learns a BGC-likeness score that recovers curated plant loci with far more complete boundaries than the unadapted model. It further claims that weak supervision derived from functional annotations can cut the fraction of primary-metabolism-like false positives roughly in half, while making predicted loci substantially more compact than a standard rule-based plant tool. If true, this would give plant researchers a label-free, learning-based way to narrow the experimental search space for specialized-metabolite pathways.

Core claim

On its own terms, the paper discovers that label-free domain adaptation plus weak supervision makes a microbe-trained Transformer a viable plant BGC detector. Concretely, on 34 curated plant BGC loci, masked-language-model adaptation on unlabeled plant sequences raises strict 100%-coverage recovery from 29.4% to 67.6%, and functional-annotation-derived weak supervision reduces a proxy primary-metabolism ratio by 48.4% and 45.2% with a paired Wilcoxon p=1.53e-5. Boundary comparison against the standard rule-based tool shows PlantBGC loci are shorter on matched regions (median length ratio 0.278; 93.8% of pairs shorter), implying lower experimental validation cost.

What carries the argument

The central object is a genome represented as an ordered sequence of conserved protein-domain tokens. The workhorse is an encoder-only Transformer architecture with a token-level scoring head: Stage 1 trains it on curated microbial BGCs with weighted binary cross-entropy; Stage 2 freezes the lower encoder blocks and the scoring head and continues pretraining the upper blocks with a masked-language-model objective on unlabeled plant sequences, aligning representations without using any plant labels; Stage 3 fine-tunes the scoring head using confidence-weighted soft negatives derived from functional annotations that flag primary-metabolism-like loci. This staged design lets the model inherit m

Load-bearing premise

The entire false-positive-control result depends on the assumption that the functional-annotation-derived 'primary-like' label reliably identifies true non-BGC loci and does not inadvertently penalize real BGCs; if that surrogate is wrong, the reported reduction in false positives is not evidence of improved specificity.

What would settle it

Take the 34 curated plant BGCs, run Stage 3, and check whether recall at 100% coverage stays at 67.6% or drops; additionally, assemble a held-out set of experimentally validated primary-metabolism gene clusters and measure how many survive as candidates. If recall collapses or primary-like loci are enriched in validated BGCs, the central claim is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Plant BGC discovery can be performed without plant labels, using microbial supervision plus unlabeled plant genomes.
  • Strict-coverage recovery of known plant loci more than doubles after label-free adaptation (29.4% to 67.6% at 100% coverage).
  • Weak supervision from functional annotations cuts the proxy primary-metabolism false-positive ratio by roughly half, with per-species consistency.
  • Predicted loci are far more compact than the rule-based comparator on matched regions (median length ratio 0.278), reducing downstream validation burden.
  • The model also generalizes across unseen microbial biosynthetic classes (leave-class-out AUC 0.979), supporting the transfer premise.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the same functional annotations generate both the training signal and the evaluation metric, an independent validation set (e.g., experimentally confirmed non-BGC regions) would be needed to rule out circularity; the paper does not supply one.
  • If the approach generalizes, the same microbe-to-eukaryote transfer strategy could be applied to fungi or other under-labeled eukaryotes, where label scarcity is equally acute.
  • A testable extension is to combine the learned score with expression or co-expression data to see whether compact boundary predictions correspond to co-regulated pathway genes.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes PlantBGC, a three-stage Transformer-based pipeline for plant biosynthetic gene cluster (BGC) discovery. Stage 1 trains an encoder-only Transformer on microbial MIBiG BGCs to score Pfam-domain tokens; Stage 2 adapts the model to plants via label-free masked language modeling on unlabeled plant CDS; Stage 3 uses GO/KEGG-derived annotations as weak supervision to down-weight primary-metabolism-like loci. The authors report strong microbial benchmark results (token-level AUC 0.988 10-fold CV, 0.979 leave-class-out), improved recovery of 34 curated plant BGCs at 100% coverage after Stage 2 (29.4% to 67.6%), a 48.4%/45.2% reduction in a GO/KEGG-derived 'primary-like' proxy ratio after Stage 3, and more compact loci than plantiSMASH on matched regions (median length ratio 0.278).

Significance. If the central claims hold, PlantBGC would be a valuable contribution: it is, to the authors' knowledge, the first ML-based plant BGC detector, and the label-free domain-adaptation idea is sensible given the scarcity of curated plant BGC labels. The microbial benchmark is solid and shows that the Transformer architecture improves over DeepBGC and a random forest in leave-class-out settings. The compactness comparison is also a useful practical outcome. The main weakness is that the Stage 3 false-positive-control claim is evaluated with the same GO/KEGG-derived labels used to construct the soft-negative training signal, making the reported reduction largely a consequence of fitting the loss. Without an independent false-positive benchmark or a post-Stage-3 recall check, the headline FP-control result is not supported.

major comments (3)
  1. [§3.4, Eqs. (7)–(15)] The Stage 3 evaluation is circular. The soft-negative targets are built from GO/KEGG evidence (Eqs. 7–12), and the reported metric, ProxyFP_reduction (Eq. 15), is the relative decrease in the fraction of predicted loci labeled primary-like by the same GO/KEGG rules. Optimizing L_soft (Eq. 12) directly minimizes the probability that predicted loci carry those annotations, so the observed 48.40%/45.20% reductions and the paired Wilcoxon p=1.53e-5 are expected consequences of fitting the model to those labels. The paper explicitly states Stage 3 'does not further aim to increase known-BGC recovery' and provides no recall check on the 34 curated BGCs after Stage 3, nor any held-out set of known non-BGC metabolic loci. As it stands, the claim that weak supervision controls false positives is not independently validated. Please provide an external evaluation—for example, recall of the curated
  2. [§3.3 and §2.4.1] The Stage 2 recovery claim rests on only 34 curated loci and fixed aggregation thresholds τ=0.5 and m=3, with no reported confidence intervals despite the text claiming 95% bootstrap CIs. It is plausible that τ and m were tuned (even informally) on this small set, in which case the 29.4%→67.6% increase at strict 100% coverage may overstate generalization. Please report bootstrap confidence intervals for the recovery differences, and provide a sensitivity analysis of τ and m on recovery and locus counts. If thresholds were chosen post hoc, this should be stated and the risk assessed.
  3. [§3.5, Fig. 4] The compactness comparison against plantiSMASH is informative but may reflect the aggregation design rather than biological boundary accuracy. PlantBGC loci are gap-free runs of CDS above τ with minimum length m=3, which by construction produces shorter, contiguous intervals; plantiSMASH uses different cluster definitions and often includes flanking regions. The paper does not report whether the compactness difference persists under varying τ and m, nor whether PlantBGC's shorter loci still achieve 100% coverage of known BGCs. Please add a threshold-sensitivity analysis for the length-ratio results and, if possible, report the number of known BGCs whose boundaries remain fully covered under the matched pairs.
minor comments (6)
  1. [Abstract and §3.5] The tool name is inconsistently capitalized: 'plantiSMASH' appears lowercase in the abstract and in several places, while the paper title and body use 'plantiSMASH'. Please standardize.
  2. [§1, Introduction] The sentence '...heterologous expression for functional validation [23], remains definitive but costly and difficult to scale' is grammatically incomplete: the subject 'workflows' is followed by an orphan comma. Please revise.
  3. [§2.5.1, Eq. (8)] The 'review' label is defined only by the catch-all 'otherwise' branch. It would help to state explicitly whether 'review' loci are excluded from soft-negative training or treated as unlabeled, since this affects the interpretation of Table 3's 'review' scale.
  4. [§3.3, Fig. 2] The text refers to 'Figure 2c' as a per-species heatmap, but the caption lists only panels (a) and (b), with (b) itself described as a heatmap. Please match captions and in-text references.
  5. [§3.2, Table 2] In Table 2, the Stage 1 row says 'ROC-AUC improves from 0.945 to 0.979 under leave-class-out', but the comparable number for the Transformer is 0.979 and for DeepBGC is 0.945. The wording 'improves from...to' is ambiguous and should identify the baseline explicitly.
  6. [§2.5.3, Eq. (13)] Eq. (13) combines L_soft with L_sup, but it is not clear whether L_sup is the original Stage 1 microbial loss or a re-computed loss on microbial data after Stage 2 adaptation. Please clarify the data and checkpoint used for L_sup in Stage 3.

Circularity Check

1 steps flagged

Stage 3's headline false-positive reduction is evaluated with the same GO/KEGG-derived labels used as its training signal; the ProxyFP drop is a training-objective echo, while Stage 1 and Stage 2 benchmarks remain independent.

specific steps
  1. fitted input called prediction [Sec. 2.5.3 (Eqs. 10–12) and Sec. 3.4 (Eq. 15, Table 4)]
    "Stage 3 fine-tunes the Stage 2 checkpoint using GO/KEGG-derived primary-like evidence as soft negatives. We assign each predicted locus a primary-like confidence from GO proxy labels and KEGG KO-based primary-metabolism signatures; loci supported by both resources receive an agreement boost (Eq. 10). The locus confidence is mapped to token-level soft targets (Eq. 11), and we optimize the confidence-weighted soft-label loss (Eq. 12) ... Let r denote the primary-like ratio among predicted candidates. ... ProxyFP_reduction(%)=100× (rStage2−rStage3)/rStage2."

    The soft-negative targets that Stage 3 optimizes (y~_t = 1 - q_l, Eq. 11; L_soft, Eq. 12) are constructed from GO/KEGG primary_likely/primary_tilt labels assigned by Eq. 8, and the reported outcome (ProxyFP_reduction, Eq. 15; Table 4) measures the change in the fraction of predicted loci classified as primary-like by the same Eq. 8 labels. Minimizing Eq. 12 directly suppresses those loci, so the 48.4%/45.2% reduction is the expected consequence of fitting the metric, not an independent test of false-positive control. No post-Stage-3 recall on the 34 curated BGCs or held-out non-BGC loci is provided; the paper states Stage 3 'does not further aim to increase known-BGC recovery.'

full rationale

Stage 1 (microbial 10-fold CV and leave-class-out) and Stage 2 (recovery of 34 curated plant BGCs after label-free MLM adaptation) are genuinely independent: the curated MIBiG/plant loci are external labels and are not used as training signals in those stages. The compactness comparison against plantiSMASH is also an external, non-circular benchmark. The circularity is confined to the Stage 3 weak-supervision claim. The GO/KEGG-derived primary-like ratio is simultaneously the target of training (via soft negatives) and the reported evaluation metric; optimizing Eq. 12 can be expected to lower Eq. 15, and the paired Wilcoxon test only confirms that the loss moved its own objective. Because the paper presents this reduction as evidence of false-positive control without an independent benchmark, this central Stage 3 claim reduces to a fit. The paper itself partially hedges by calling this a 'proxy' and by not claiming recovery gains for Stage 3, but the FP-control interpretation is still load-bearing in the abstract and intro. Overall score 7: one major circular evaluation embedded in an otherwise self-contained pipeline; not a fully circular derivation, but the Stage 3 headline result is forced by construction.

Axiom & Free-Parameter Ledger

9 free parameters · 7 axioms · 0 invented entities

The paper adds no new physical or biological entities. Its load-bearing assumptions are representational (Pfam tokens), transfer (microbes to plants), and evaluative (GO/KEGG proxy as gold standard for false positives). The last is the least supported.

free parameters (9)
  • α (positive-class weight in BCE) = not reported
    Eq. 4; set based on class imbalance; affects precision/recall trade-off.
  • τ (CDS score threshold) = 0.5
    Section 2.4.1; default threshold for selecting high-scoring CDS; not shown to be tuned on held-out plant loci.
  • m (minimum locus length) = 3
    Section 2.4.1; minimum consecutive CDS count; suppresses spurious short runs.
  • gap tolerance = 0
    Section 2.4.1; strict consecutiveness.
  • ρ (masking rate) = 0.15
    Section 2.3.2; standard MLM masking.
  • γ (agreement boost) = 0.5
    Eq. 10; fuses GO and KEGG confidences.
  • soft-label confidences for primary_likely/primary_tilt = 0.8 / 0.5
    Section 2.5.2; hand-chosen confidence mapping.
  • λ (soft-loss weight) = not reported
    Eq. 13; balances Lsup and Lsoft; unknown in text.
  • Transformer hyperparameters (L, d_model, heads, dropout) = not reported
    Section 2.3; not specified in text, needed for reproduction.
axioms (7)
  • domain assumption Pfam-domain tokens plus Pfam2vec embeddings are a sufficient representation of BGC-likeness for both microbes and plants
    Section 2.2; all inputs are ordered Pfam tokens; if domain composition misses BGC signal, pipeline fails.
  • domain assumption Microbial BGC supervision (MIBiG) transfers to plants after MLM adaptation without plant labels
    Section 2.3.2; central premise of the transfer approach.
  • domain assumption GO/KEGG term sets for primary/secondary metabolism are correctly defined and comprehensive
    Section 2.5.1; fixes the proxy labels; if the GO/KEGG assignments are wrong, the weak supervision is misdirected.
  • domain assumption The 34 curated plant BGC intervals are correct and complete on the reference assemblies
    Section 3.1; the recovery metric depends on these intervals.
  • domain assumption GeneSwap-generated negatives are valid non-BGC representatives
    Section 3.1; used for Stage 1 training/evaluation.
  • domain assumption MLM with frozen lower blocks and scoring head preserves the learned decision function
    Section 2.3.2; if adaptation destroys the boundary, Stage 2 gains would not be meaningful.
  • standard math Weighted BCE and soft-label BCE objectives are appropriate for the task
    Eqs. 4, 12; standard objectives.

pith-pipeline@v1.3.0-daily-deepseek · 14246 in / 17803 out tokens · 174233 ms · 2026-08-01T16:11:06.342539+00:00 · methodology

0 comments
read the original abstract

Plant biosynthetic gene clusters (BGCs) encode specialized-metabolite pathways, yet curated plant BGC labels remain scarce, hindering supervised discovery at genome scale. Existing plant BGC mining tools are largely signature- and rule-driven and do not fully leverage recent advances in contextual representation learning for modeling long-range domain context and controlling false positives under strong domain shift. We seek an AI-assisted workflow that narrows experimental search space by transferring supervision from well-annotated microbial BGCs to plant genomes. We present PlantBGC, representing genomes as ordered Pfam-domain sequences and learning BGC-likeness with an encoder-only Transformer trained on MIBiG microbial BGCs and adapted to plants via label-free masked language modeling. On microbial benchmarks, PlantBGC achieves token-level AUC = 0.988 (10-fold CV) and 0.979 (leave-class-out). On plants, adaptation improves known-BGC recovery on n = 34 curated loci under strict 100% coverage, increasing recovery from 29.4% to 67.6% and indicating more complete boundaries. GO/KEGG-derived weak supervision reduces proxy primary-like ratio by 48.40% (GO) and 45.20% (KEGG), with consistent per-species reductions (paired Wilcoxon p = 1.53e-5). Compared to plantiSMASH, PlantBGC yields more compact loci on matched regions (median length ratio = 0.278; 93.8% of pairs are shorter).

Figures

Figures reproduced from arXiv: 2607.27258 by Nidhi Grover, Ning Sui, Yuhan Zhao, Zhishan Guo.

Figure 1
Figure 1. Figure 1: Overview of PlantBGC with label transfer and weak supervision. (A) Input construction for microbial and plant data. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Stage 2 unlabeled plant adaptation improves known-BGC recovery. (a) Recovery of [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Stage 3 reduces proxy primary-metabolism burden [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Bidirectional compactness comparison between PlantBGC and plantiSMASH. Pairs are defined by best matches with [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

32 extracted references · 14 canonical work pages

  1. [1]

    Adrio and Arnold L

    Jose L. Adrio and Arnold L. Demain. 2006. Genetic improvement of processes yielding microbial products.FEMS Microbiology Reviews30, 2 (2006), 187–214. doi:10.1111/j.1574-6976.2005.00009.x

  2. [2]

    Altschul, Warren Gish, Webb Miller, Eugene W

    Stephen F. Altschul, Warren Gish, Webb Miller, Eugene W. Myers, and David J. Lipman. 1990. Basic local alignment search tool.Journal of Molecular Biology 215, 3 (1990), 403–410. doi:10.1016/S0022-2836(05)80360-2

  3. [3]

    Ball, Judith A

    Michael Ashburner, Catherine A. Ball, Judith A. Blake, David Botstein, Heather Butler, J. Michael Cherry, Allan P. Davis, Kara Dolinski, Selina S. Dwight, Janan T. Eppig, et al. 2000. Gene ontology: tool for the unification of biology.Nature Genetics25, 1 (2000), 25–29. doi:10.1038/75556

  4. [4]

    Bauman, Keelie S

    Katherine D. Bauman, Keelie S. Butler, Bradley S. Moore, and Jonathan R. Chekan

  5. [5]

    Augustijn, Zachary L

    Kai Blin, Simon Shaw, Hannah E. Augustijn, Zachary L. Reitz, Friederike Bier- mann, Mohammad Alanjary, Artem Fetter, Barbara R. Terlouw, William W. Met- calf, Eric J. N. Helfrich, Gilles P. van Wezel, Marnix H. Medema, and Tilmann Weber. 2023. antiSMASH 7.0: new and improved predictions for detection, reg- ulation, chemical structures and visualisation.Nu...

  6. [6]

    Leo Breiman. 2001. Random forests.Machine Learning45 (2001), 5–32. doi:10. 1023/A:1010933404324

  7. [7]

    Carroll, Martin Larralde, Jonas S

    Laura M. Carroll, Martin Larralde, Jonas S. Fleck, et al. 2021. Accurate de novo identification of biosynthetic gene clusters with GECCO.bioRxiv(2021). doi:10. 1101/2021.05.03.442509

  8. [8]

    Medema, Jan Claesen, Kenji Kurita, Laura C

    Peter Cimermancic, Marnix H. Medema, Jan Claesen, Kenji Kurita, Laura C. Wieland Brown, Konstantinos Mavrommatis, Amrita Pati, Paul A. Godfrey, Michael Koehrsen, Jon Clardy, et al . 2014. Insights into secondary metabo- lism from a global analysis of prokaryotic biosynthetic gene clusters.Cell158, 2 (2014), 412–421. doi:10.1016/j.cell.2014.06.034

  9. [9]

    Dias, Sylvia Urban, and Ute Roessner

    Daniel A. Dias, Sylvia Urban, and Ute Roessner. 2012. A historical overview of natural products in drug discovery.Metabolites2, 2 (2012), 303–336. doi:10.3390/ metabo2020303

  10. [10]

    Sean R. Eddy. 2011. Accelerated profile HMM searches.PLoS Computational Biology7, 10 (2011), e1002195. doi:10.1371/journal.pcbi.1002195

  11. [11]

    Eddy, et al

    Sara El-Gebali, Jaina Mistry, Alex Bateman, Sean R. Eddy, et al. 2019. The Pfam protein families database in 2019.Nucleic Acids Research47, D1 (2019), D427– D432. doi:10.1093/nar/gky995

  12. [12]

    Hannigan, Daniel Prihoda, Adam Palicka, Jan Soukup, Ondrej Klem- pir, Leena Rampula, et al

    Gavin D. Hannigan, Daniel Prihoda, Adam Palicka, Jan Soukup, Ondrej Klem- pir, Leena Rampula, et al. 2019. A deep learning genome-mining strategy for biosynthetic gene cluster prediction.Nucleic Acids Research47, 18 (2019), e110. doi:10.1093/nar/gkz654

  13. [13]

    Jaewook Hwang, Jonathan Kirshner, Daniel A. R. Ramey Deschênes, Matthew B. Richardson, Steven J. Fleck, others, and Yang Qu. 2025. Ancient gene clusters govern the initiation of monoterpenoid indole alkaloid biosynthesis and C3 stereochemistry inversion.Nature Communications16 (2025), 10495. doi:10.1038/ s41467-025-65543-z

  14. [14]

    LoCascio, Miriam Land, Frank W

    Doug Hyatt, Gwo-Liang Chen, Philip F. LoCascio, Miriam Land, Frank W. Larimer, and Loren J. Hauser. 2010. Prodigal: prokaryotic gene recognition and translation initiation site identification.BMC Bioinformatics11 (2010), 119. doi:10.1186/1471- 2105-11-119

  15. [15]

    Minoru Kanehisa and Susumu Goto. 2000. KEGG: Kyoto Encyclopedia of Genes and Genomes.Nucleic Acids Research28, 1 (2000), 27–30. doi:10.1093/nar/28.1.27

  16. [16]

    Kautsar, Hernando G

    Satria A. Kautsar, Hernando G. Suarez Duran, Kai Blin, Anne Osbourn, and Marnix H. Medema. 2017. plantiSMASH: automated identification, annotation and expression analysis of plant biosynthetic gene clusters.Nucleic Acids Research 45, W1 (2017), W55–W63. doi:10.1093/nar/gkx305

  17. [17]

    Tomoki Kawano, Taro Shiraishi, Tomohisa Kuzuyama, and Masayuki Umemura

  18. [18]

    Mingyang Liu, Yun Li, and Hongzhe Li. 2022. Deep learning to predict the biosynthetic gene clusters in bacterial genomes.Journal of Molecular Biology 434, 14 (2022), 167463. doi:10.1016/j.jmb.2022.167463

  19. [19]

    Madival, Dwijesh Chandra Mishra, Krishna Kumar Chaturvedi, Neeraj Budhlakoti, Mohammad Samir Farooqi, Sudhir Srivastava, Anu Sharma, Shivadarshan S

    Sharanbasappa D. Madival, Dwijesh Chandra Mishra, Krishna Kumar Chaturvedi, Neeraj Budhlakoti, Mohammad Samir Farooqi, Sudhir Srivastava, Anu Sharma, Shivadarshan S. Jirli, Alka Arora, Girish K. Jha, and Shesh N. Rai. 2025. RF- BGCpred: A random forest based tool for prediction of biosynthetic gene clusters. Biochimica et Biophysica Acta (BBA) - General S...

  20. [20]

    Medema, Ralf Kottmann, Raphael Yilmaz, Markus Cummings, John B

    Marnix H. Medema, Ralf Kottmann, Raphael Yilmaz, Markus Cummings, John B. Biggins, Kai Blin, et al. 2015. Minimum information about a biosynthetic gene cluster.Nature Chemical Biology11 (2015), 625–631. doi:10.1038/nchembio.1890

  21. [21]

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space.arXiv(2013). https://arxiv. org/abs/1301.3781

  22. [22]

    Newman and Gordon M

    David J. Newman and Gordon M. Cragg. 2016. Natural products as sources of new drugs from 1981 to 2014.Journal of Natural Products79, 3 (2016), 629–661. doi:10.1021/acs.jnatprod.5b01055

  23. [23]

    Yaodong Ning, Yao Xu, Binghua Jiao, and Xiaoling Lu. 2022. Application of Gene Knockout and Heterologous Expression Strategy in Fungal Secondary Metabo- lites Biosynthesis.Marine Drugs20, 11 (2022), 705. doi:10.3390/md20110705

  24. [24]

    Reeves, R

    Andrew R. Reeves, R. Samuel English, J. S. Lampel, David A. Post, and Thomas J. Vanden Boom. 1999. Transcriptional organization of the erythromycin biosyn- thetic gene cluster of Saccharopolyspora erythraea.Journal of Bacteriology181, 22 (1999), 7098–7106. doi:10.1128/JB.181.22.7098-7106.1999

  25. [25]

    Carolina Rios-Martinez, Nitin Bhattacharya, A. P. Amini, Lucas Crawford, and Kevin K. Yang. 2023. Deep self-supervised learning for biosynthetic gene cluster detection and product classification.PLOS Computational Biology19, 5 (2023), e1011162. doi:10.1371/journal.pcbi.1011162

  26. [26]

    Chavali, Ricardo Nilo-Poyanco, Thomas Bernard, Daniel Kahn, and Seung Y

    Pascal Schläpfer, Peifen Zhang, Chuan Wang, Taehyong Kim, Michael Banf, Lee Chae, Kate Dreher, Arvind K. Chavali, Ricardo Nilo-Poyanco, Thomas Bernard, Daniel Kahn, and Seung Y. Rhee. 2017. Genome-Wide Prediction of Metabolic Enzymes, Pathways, and Gene Clusters in Plants.Plant Physiology173, 4 (2017), 2041–2059. doi:10.1104/pp.16.01942

  27. [27]

    Skinnider, Chris A

    Michael A. Skinnider, Chris A. Dejong, Philip N. Rees, Chad W. Johnston, Haoxin Li, Andrew L. H. Webster, Morgan A. Wyatt, and Nathan A. Magarvey. 2015. Genomes to natural products PRediction Informatics for Secondary Metabolomes (PRISM).Nucleic Acids Research43, 20 (2015), 9645–9662. doi:10.1093/nar/gkv1012

  28. [28]

    Nadine Töpfer, Lisa-Maria Fuchs, and Asaph Aharoni. 2017. The PhytoClust tool for metabolic gene clusters discovery in plant genomes.Nucleic Acids Research 45, 12 (2017), 7049–7063. doi:10.1093/nar/gkx404

  29. [29]

    Xingyao Xiong, Yan Li, Yanjun Zeng, et al. 2021. The Taxus genome provides insights into paclitaxel biosynthesis.Nature Plants7, 8 (2021), 1026–1036. doi:10. 1038/s41477-021-00963-5

  30. [30]

    Zdouc, Kai Blin, Nico L

    Mitja M. Zdouc, Kai Blin, Nico L. L. Louwen, Jorge Navarro, et al. 2025. MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration. Nucleic Acids Research53, D1 (2025), D678–D690. doi:10.1093/nar/gkae1115

  31. [2021]

    doi:10.1039/d1np00025h

    Genome mining methods to discover bioactive natural products.Natural Product Reports38, 11 (2021), 2100–2129. doi:10.1039/d1np00025h

  32. [2025]

    doi:10.1101/ 2025.06.02.657346

    A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products.bioRxiv(2025). doi:10.1101/ 2025.06.02.657346