REVIEW 3 major objections 6 minor 32 references
A Transformer originally trained on microbial biosynthetic gene clusters can, after label-free adaptation on unlabeled plant genomes and weak supervision from functional annotations, discover plant BGCs without any curated plant labels, rec
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 16:11 UTC pith:X2DIFFDZ
load-bearing objection First learning-based plant BGC discovery pipeline worth taking seriously, but the headline false-positive reduction is circular until tested on labels other than the ones that trained it. the 3 major comments →
PlantBGC: Transformer for Plant BGC Discovery via Label-Free Domain Adaptation and Weak Supervision
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On its own terms, the paper discovers that label-free domain adaptation plus weak supervision makes a microbe-trained Transformer a viable plant BGC detector. Concretely, on 34 curated plant BGC loci, masked-language-model adaptation on unlabeled plant sequences raises strict 100%-coverage recovery from 29.4% to 67.6%, and functional-annotation-derived weak supervision reduces a proxy primary-metabolism ratio by 48.4% and 45.2% with a paired Wilcoxon p=1.53e-5. Boundary comparison against the standard rule-based tool shows PlantBGC loci are shorter on matched regions (median length ratio 0.278; 93.8% of pairs shorter), implying lower experimental validation cost.
What carries the argument
The central object is a genome represented as an ordered sequence of conserved protein-domain tokens. The workhorse is an encoder-only Transformer architecture with a token-level scoring head: Stage 1 trains it on curated microbial BGCs with weighted binary cross-entropy; Stage 2 freezes the lower encoder blocks and the scoring head and continues pretraining the upper blocks with a masked-language-model objective on unlabeled plant sequences, aligning representations without using any plant labels; Stage 3 fine-tunes the scoring head using confidence-weighted soft negatives derived from functional annotations that flag primary-metabolism-like loci. This staged design lets the model inherit m
Load-bearing premise
The entire false-positive-control result depends on the assumption that the functional-annotation-derived 'primary-like' label reliably identifies true non-BGC loci and does not inadvertently penalize real BGCs; if that surrogate is wrong, the reported reduction in false positives is not evidence of improved specificity.
What would settle it
Take the 34 curated plant BGCs, run Stage 3, and check whether recall at 100% coverage stays at 67.6% or drops; additionally, assemble a held-out set of experimentally validated primary-metabolism gene clusters and measure how many survive as candidates. If recall collapses or primary-like loci are enriched in validated BGCs, the central claim is falsified.
If this is right
- Plant BGC discovery can be performed without plant labels, using microbial supervision plus unlabeled plant genomes.
- Strict-coverage recovery of known plant loci more than doubles after label-free adaptation (29.4% to 67.6% at 100% coverage).
- Weak supervision from functional annotations cuts the proxy primary-metabolism false-positive ratio by roughly half, with per-species consistency.
- Predicted loci are far more compact than the rule-based comparator on matched regions (median length ratio 0.278), reducing downstream validation burden.
- The model also generalizes across unseen microbial biosynthetic classes (leave-class-out AUC 0.979), supporting the transfer premise.
Where Pith is reading between the lines
- Because the same functional annotations generate both the training signal and the evaluation metric, an independent validation set (e.g., experimentally confirmed non-BGC regions) would be needed to rule out circularity; the paper does not supply one.
- If the approach generalizes, the same microbe-to-eukaryote transfer strategy could be applied to fungi or other under-labeled eukaryotes, where label scarcity is equally acute.
- A testable extension is to combine the learned score with expression or co-expression data to see whether compact boundary predictions correspond to co-regulated pathway genes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PlantBGC, a three-stage Transformer-based pipeline for plant biosynthetic gene cluster (BGC) discovery. Stage 1 trains an encoder-only Transformer on microbial MIBiG BGCs to score Pfam-domain tokens; Stage 2 adapts the model to plants via label-free masked language modeling on unlabeled plant CDS; Stage 3 uses GO/KEGG-derived annotations as weak supervision to down-weight primary-metabolism-like loci. The authors report strong microbial benchmark results (token-level AUC 0.988 10-fold CV, 0.979 leave-class-out), improved recovery of 34 curated plant BGCs at 100% coverage after Stage 2 (29.4% to 67.6%), a 48.4%/45.2% reduction in a GO/KEGG-derived 'primary-like' proxy ratio after Stage 3, and more compact loci than plantiSMASH on matched regions (median length ratio 0.278).
Significance. If the central claims hold, PlantBGC would be a valuable contribution: it is, to the authors' knowledge, the first ML-based plant BGC detector, and the label-free domain-adaptation idea is sensible given the scarcity of curated plant BGC labels. The microbial benchmark is solid and shows that the Transformer architecture improves over DeepBGC and a random forest in leave-class-out settings. The compactness comparison is also a useful practical outcome. The main weakness is that the Stage 3 false-positive-control claim is evaluated with the same GO/KEGG-derived labels used to construct the soft-negative training signal, making the reported reduction largely a consequence of fitting the loss. Without an independent false-positive benchmark or a post-Stage-3 recall check, the headline FP-control result is not supported.
major comments (3)
- [§3.4, Eqs. (7)–(15)] The Stage 3 evaluation is circular. The soft-negative targets are built from GO/KEGG evidence (Eqs. 7–12), and the reported metric, ProxyFP_reduction (Eq. 15), is the relative decrease in the fraction of predicted loci labeled primary-like by the same GO/KEGG rules. Optimizing L_soft (Eq. 12) directly minimizes the probability that predicted loci carry those annotations, so the observed 48.40%/45.20% reductions and the paired Wilcoxon p=1.53e-5 are expected consequences of fitting the model to those labels. The paper explicitly states Stage 3 'does not further aim to increase known-BGC recovery' and provides no recall check on the 34 curated BGCs after Stage 3, nor any held-out set of known non-BGC metabolic loci. As it stands, the claim that weak supervision controls false positives is not independently validated. Please provide an external evaluation—for example, recall of the curated
- [§3.3 and §2.4.1] The Stage 2 recovery claim rests on only 34 curated loci and fixed aggregation thresholds τ=0.5 and m=3, with no reported confidence intervals despite the text claiming 95% bootstrap CIs. It is plausible that τ and m were tuned (even informally) on this small set, in which case the 29.4%→67.6% increase at strict 100% coverage may overstate generalization. Please report bootstrap confidence intervals for the recovery differences, and provide a sensitivity analysis of τ and m on recovery and locus counts. If thresholds were chosen post hoc, this should be stated and the risk assessed.
- [§3.5, Fig. 4] The compactness comparison against plantiSMASH is informative but may reflect the aggregation design rather than biological boundary accuracy. PlantBGC loci are gap-free runs of CDS above τ with minimum length m=3, which by construction produces shorter, contiguous intervals; plantiSMASH uses different cluster definitions and often includes flanking regions. The paper does not report whether the compactness difference persists under varying τ and m, nor whether PlantBGC's shorter loci still achieve 100% coverage of known BGCs. Please add a threshold-sensitivity analysis for the length-ratio results and, if possible, report the number of known BGCs whose boundaries remain fully covered under the matched pairs.
minor comments (6)
- [Abstract and §3.5] The tool name is inconsistently capitalized: 'plantiSMASH' appears lowercase in the abstract and in several places, while the paper title and body use 'plantiSMASH'. Please standardize.
- [§1, Introduction] The sentence '...heterologous expression for functional validation [23], remains definitive but costly and difficult to scale' is grammatically incomplete: the subject 'workflows' is followed by an orphan comma. Please revise.
- [§2.5.1, Eq. (8)] The 'review' label is defined only by the catch-all 'otherwise' branch. It would help to state explicitly whether 'review' loci are excluded from soft-negative training or treated as unlabeled, since this affects the interpretation of Table 3's 'review' scale.
- [§3.3, Fig. 2] The text refers to 'Figure 2c' as a per-species heatmap, but the caption lists only panels (a) and (b), with (b) itself described as a heatmap. Please match captions and in-text references.
- [§3.2, Table 2] In Table 2, the Stage 1 row says 'ROC-AUC improves from 0.945 to 0.979 under leave-class-out', but the comparable number for the Transformer is 0.979 and for DeepBGC is 0.945. The wording 'improves from...to' is ambiguous and should identify the baseline explicitly.
- [§2.5.3, Eq. (13)] Eq. (13) combines L_soft with L_sup, but it is not clear whether L_sup is the original Stage 1 microbial loss or a re-computed loss on microbial data after Stage 2 adaptation. Please clarify the data and checkpoint used for L_sup in Stage 3.
Circularity Check
Stage 3's headline false-positive reduction is evaluated with the same GO/KEGG-derived labels used as its training signal; the ProxyFP drop is a training-objective echo, while Stage 1 and Stage 2 benchmarks remain independent.
specific steps
-
fitted input called prediction
[Sec. 2.5.3 (Eqs. 10–12) and Sec. 3.4 (Eq. 15, Table 4)]
"Stage 3 fine-tunes the Stage 2 checkpoint using GO/KEGG-derived primary-like evidence as soft negatives. We assign each predicted locus a primary-like confidence from GO proxy labels and KEGG KO-based primary-metabolism signatures; loci supported by both resources receive an agreement boost (Eq. 10). The locus confidence is mapped to token-level soft targets (Eq. 11), and we optimize the confidence-weighted soft-label loss (Eq. 12) ... Let r denote the primary-like ratio among predicted candidates. ... ProxyFP_reduction(%)=100× (rStage2−rStage3)/rStage2."
The soft-negative targets that Stage 3 optimizes (y~_t = 1 - q_l, Eq. 11; L_soft, Eq. 12) are constructed from GO/KEGG primary_likely/primary_tilt labels assigned by Eq. 8, and the reported outcome (ProxyFP_reduction, Eq. 15; Table 4) measures the change in the fraction of predicted loci classified as primary-like by the same Eq. 8 labels. Minimizing Eq. 12 directly suppresses those loci, so the 48.4%/45.2% reduction is the expected consequence of fitting the metric, not an independent test of false-positive control. No post-Stage-3 recall on the 34 curated BGCs or held-out non-BGC loci is provided; the paper states Stage 3 'does not further aim to increase known-BGC recovery.'
full rationale
Stage 1 (microbial 10-fold CV and leave-class-out) and Stage 2 (recovery of 34 curated plant BGCs after label-free MLM adaptation) are genuinely independent: the curated MIBiG/plant loci are external labels and are not used as training signals in those stages. The compactness comparison against plantiSMASH is also an external, non-circular benchmark. The circularity is confined to the Stage 3 weak-supervision claim. The GO/KEGG-derived primary-like ratio is simultaneously the target of training (via soft negatives) and the reported evaluation metric; optimizing Eq. 12 can be expected to lower Eq. 15, and the paired Wilcoxon test only confirms that the loss moved its own objective. Because the paper presents this reduction as evidence of false-positive control without an independent benchmark, this central Stage 3 claim reduces to a fit. The paper itself partially hedges by calling this a 'proxy' and by not claiming recovery gains for Stage 3, but the FP-control interpretation is still load-bearing in the abstract and intro. Overall score 7: one major circular evaluation embedded in an otherwise self-contained pipeline; not a fully circular derivation, but the Stage 3 headline result is forced by construction.
Axiom & Free-Parameter Ledger
free parameters (9)
- α (positive-class weight in BCE) =
not reported
- τ (CDS score threshold) =
0.5
- m (minimum locus length) =
3
- gap tolerance =
0
- ρ (masking rate) =
0.15
- γ (agreement boost) =
0.5
- soft-label confidences for primary_likely/primary_tilt =
0.8 / 0.5
- λ (soft-loss weight) =
not reported
- Transformer hyperparameters (L, d_model, heads, dropout) =
not reported
axioms (7)
- domain assumption Pfam-domain tokens plus Pfam2vec embeddings are a sufficient representation of BGC-likeness for both microbes and plants
- domain assumption Microbial BGC supervision (MIBiG) transfers to plants after MLM adaptation without plant labels
- domain assumption GO/KEGG term sets for primary/secondary metabolism are correctly defined and comprehensive
- domain assumption The 34 curated plant BGC intervals are correct and complete on the reference assemblies
- domain assumption GeneSwap-generated negatives are valid non-BGC representatives
- domain assumption MLM with frozen lower blocks and scoring head preserves the learned decision function
- standard math Weighted BCE and soft-label BCE objectives are appropriate for the task
read the original abstract
Plant biosynthetic gene clusters (BGCs) encode specialized-metabolite pathways, yet curated plant BGC labels remain scarce, hindering supervised discovery at genome scale. Existing plant BGC mining tools are largely signature- and rule-driven and do not fully leverage recent advances in contextual representation learning for modeling long-range domain context and controlling false positives under strong domain shift. We seek an AI-assisted workflow that narrows experimental search space by transferring supervision from well-annotated microbial BGCs to plant genomes. We present PlantBGC, representing genomes as ordered Pfam-domain sequences and learning BGC-likeness with an encoder-only Transformer trained on MIBiG microbial BGCs and adapted to plants via label-free masked language modeling. On microbial benchmarks, PlantBGC achieves token-level AUC = 0.988 (10-fold CV) and 0.979 (leave-class-out). On plants, adaptation improves known-BGC recovery on n = 34 curated loci under strict 100% coverage, increasing recovery from 29.4% to 67.6% and indicating more complete boundaries. GO/KEGG-derived weak supervision reduces proxy primary-like ratio by 48.40% (GO) and 45.20% (KEGG), with consistent per-species reductions (paired Wilcoxon p = 1.53e-5). Compared to plantiSMASH, PlantBGC yields more compact loci on matched regions (median length ratio = 0.278; 93.8% of pairs are shorter).
Figures
Reference graph
Works this paper leans on
-
[1]
Jose L. Adrio and Arnold L. Demain. 2006. Genetic improvement of processes yielding microbial products.FEMS Microbiology Reviews30, 2 (2006), 187–214. doi:10.1111/j.1574-6976.2005.00009.x
arXiv 2006
-
[2]
Altschul, Warren Gish, Webb Miller, Eugene W
Stephen F. Altschul, Warren Gish, Webb Miller, Eugene W. Myers, and David J. Lipman. 1990. Basic local alignment search tool.Journal of Molecular Biology 215, 3 (1990), 403–410. doi:10.1016/S0022-2836(05)80360-2
-
[3]
Michael Ashburner, Catherine A. Ball, Judith A. Blake, David Botstein, Heather Butler, J. Michael Cherry, Allan P. Davis, Kara Dolinski, Selina S. Dwight, Janan T. Eppig, et al. 2000. Gene ontology: tool for the unification of biology.Nature Genetics25, 1 (2000), 25–29. doi:10.1038/75556
doi:10.1038/75556 2000
-
[4]
Bauman, Keelie S
Katherine D. Bauman, Keelie S. Butler, Bradley S. Moore, and Jonathan R. Chekan
-
[5]
Kai Blin, Simon Shaw, Hannah E. Augustijn, Zachary L. Reitz, Friederike Bier- mann, Mohammad Alanjary, Artem Fetter, Barbara R. Terlouw, William W. Met- calf, Eric J. N. Helfrich, Gilles P. van Wezel, Marnix H. Medema, and Tilmann Weber. 2023. antiSMASH 7.0: new and improved predictions for detection, reg- ulation, chemical structures and visualisation.Nu...
-
[6]
Leo Breiman. 2001. Random forests.Machine Learning45 (2001), 5–32. doi:10. 1023/A:1010933404324
2001
-
[7]
Carroll, Martin Larralde, Jonas S
Laura M. Carroll, Martin Larralde, Jonas S. Fleck, et al. 2021. Accurate de novo identification of biosynthetic gene clusters with GECCO.bioRxiv(2021). doi:10. 1101/2021.05.03.442509
2021
-
[8]
Medema, Jan Claesen, Kenji Kurita, Laura C
Peter Cimermancic, Marnix H. Medema, Jan Claesen, Kenji Kurita, Laura C. Wieland Brown, Konstantinos Mavrommatis, Amrita Pati, Paul A. Godfrey, Michael Koehrsen, Jon Clardy, et al . 2014. Insights into secondary metabo- lism from a global analysis of prokaryotic biosynthetic gene clusters.Cell158, 2 (2014), 412–421. doi:10.1016/j.cell.2014.06.034
-
[9]
Dias, Sylvia Urban, and Ute Roessner
Daniel A. Dias, Sylvia Urban, and Ute Roessner. 2012. A historical overview of natural products in drug discovery.Metabolites2, 2 (2012), 303–336. doi:10.3390/ metabo2020303
2012
-
[10]
Sean R. Eddy. 2011. Accelerated profile HMM searches.PLoS Computational Biology7, 10 (2011), e1002195. doi:10.1371/journal.pcbi.1002195
-
[11]
Sara El-Gebali, Jaina Mistry, Alex Bateman, Sean R. Eddy, et al. 2019. The Pfam protein families database in 2019.Nucleic Acids Research47, D1 (2019), D427– D432. doi:10.1093/nar/gky995
-
[12]
Hannigan, Daniel Prihoda, Adam Palicka, Jan Soukup, Ondrej Klem- pir, Leena Rampula, et al
Gavin D. Hannigan, Daniel Prihoda, Adam Palicka, Jan Soukup, Ondrej Klem- pir, Leena Rampula, et al. 2019. A deep learning genome-mining strategy for biosynthetic gene cluster prediction.Nucleic Acids Research47, 18 (2019), e110. doi:10.1093/nar/gkz654
-
[13]
Jaewook Hwang, Jonathan Kirshner, Daniel A. R. Ramey Deschênes, Matthew B. Richardson, Steven J. Fleck, others, and Yang Qu. 2025. Ancient gene clusters govern the initiation of monoterpenoid indole alkaloid biosynthesis and C3 stereochemistry inversion.Nature Communications16 (2025), 10495. doi:10.1038/ s41467-025-65543-z
2025
-
[14]
LoCascio, Miriam Land, Frank W
Doug Hyatt, Gwo-Liang Chen, Philip F. LoCascio, Miriam Land, Frank W. Larimer, and Loren J. Hauser. 2010. Prodigal: prokaryotic gene recognition and translation initiation site identification.BMC Bioinformatics11 (2010), 119. doi:10.1186/1471- 2105-11-119
doi:10.1186/1471- 2010
-
[15]
Minoru Kanehisa and Susumu Goto. 2000. KEGG: Kyoto Encyclopedia of Genes and Genomes.Nucleic Acids Research28, 1 (2000), 27–30. doi:10.1093/nar/28.1.27
-
[16]
Satria A. Kautsar, Hernando G. Suarez Duran, Kai Blin, Anne Osbourn, and Marnix H. Medema. 2017. plantiSMASH: automated identification, annotation and expression analysis of plant biosynthetic gene clusters.Nucleic Acids Research 45, W1 (2017), W55–W63. doi:10.1093/nar/gkx305
-
[17]
Tomoki Kawano, Taro Shiraishi, Tomohisa Kuzuyama, and Masayuki Umemura
-
[18]
Mingyang Liu, Yun Li, and Hongzhe Li. 2022. Deep learning to predict the biosynthetic gene clusters in bacterial genomes.Journal of Molecular Biology 434, 14 (2022), 167463. doi:10.1016/j.jmb.2022.167463
arXiv 2022
-
[19]
Sharanbasappa D. Madival, Dwijesh Chandra Mishra, Krishna Kumar Chaturvedi, Neeraj Budhlakoti, Mohammad Samir Farooqi, Sudhir Srivastava, Anu Sharma, Shivadarshan S. Jirli, Alka Arora, Girish K. Jha, and Shesh N. Rai. 2025. RF- BGCpred: A random forest based tool for prediction of biosynthetic gene clusters. Biochimica et Biophysica Acta (BBA) - General S...
arXiv 2025
-
[20]
Medema, Ralf Kottmann, Raphael Yilmaz, Markus Cummings, John B
Marnix H. Medema, Ralf Kottmann, Raphael Yilmaz, Markus Cummings, John B. Biggins, Kai Blin, et al. 2015. Minimum information about a biosynthetic gene cluster.Nature Chemical Biology11 (2015), 625–631. doi:10.1038/nchembio.1890
-
[21]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space.arXiv(2013). https://arxiv. org/abs/1301.3781
Pith/arXiv arXiv 2013
-
[22]
David J. Newman and Gordon M. Cragg. 2016. Natural products as sources of new drugs from 1981 to 2014.Journal of Natural Products79, 3 (2016), 629–661. doi:10.1021/acs.jnatprod.5b01055
-
[23]
Yaodong Ning, Yao Xu, Binghua Jiao, and Xiaoling Lu. 2022. Application of Gene Knockout and Heterologous Expression Strategy in Fungal Secondary Metabo- lites Biosynthesis.Marine Drugs20, 11 (2022), 705. doi:10.3390/md20110705
-
[24]
Andrew R. Reeves, R. Samuel English, J. S. Lampel, David A. Post, and Thomas J. Vanden Boom. 1999. Transcriptional organization of the erythromycin biosyn- thetic gene cluster of Saccharopolyspora erythraea.Journal of Bacteriology181, 22 (1999), 7098–7106. doi:10.1128/JB.181.22.7098-7106.1999
arXiv 1999
-
[25]
Carolina Rios-Martinez, Nitin Bhattacharya, A. P. Amini, Lucas Crawford, and Kevin K. Yang. 2023. Deep self-supervised learning for biosynthetic gene cluster detection and product classification.PLOS Computational Biology19, 5 (2023), e1011162. doi:10.1371/journal.pcbi.1011162
-
[26]
Chavali, Ricardo Nilo-Poyanco, Thomas Bernard, Daniel Kahn, and Seung Y
Pascal Schläpfer, Peifen Zhang, Chuan Wang, Taehyong Kim, Michael Banf, Lee Chae, Kate Dreher, Arvind K. Chavali, Ricardo Nilo-Poyanco, Thomas Bernard, Daniel Kahn, and Seung Y. Rhee. 2017. Genome-Wide Prediction of Metabolic Enzymes, Pathways, and Gene Clusters in Plants.Plant Physiology173, 4 (2017), 2041–2059. doi:10.1104/pp.16.01942
-
[27]
Michael A. Skinnider, Chris A. Dejong, Philip N. Rees, Chad W. Johnston, Haoxin Li, Andrew L. H. Webster, Morgan A. Wyatt, and Nathan A. Magarvey. 2015. Genomes to natural products PRediction Informatics for Secondary Metabolomes (PRISM).Nucleic Acids Research43, 20 (2015), 9645–9662. doi:10.1093/nar/gkv1012
-
[28]
Nadine Töpfer, Lisa-Maria Fuchs, and Asaph Aharoni. 2017. The PhytoClust tool for metabolic gene clusters discovery in plant genomes.Nucleic Acids Research 45, 12 (2017), 7049–7063. doi:10.1093/nar/gkx404
-
[29]
Xingyao Xiong, Yan Li, Yanjun Zeng, et al. 2021. The Taxus genome provides insights into paclitaxel biosynthesis.Nature Plants7, 8 (2021), 1026–1036. doi:10. 1038/s41477-021-00963-5
2021
-
[30]
Mitja M. Zdouc, Kai Blin, Nico L. L. Louwen, Jorge Navarro, et al. 2025. MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration. Nucleic Acids Research53, D1 (2025), D678–D690. doi:10.1093/nar/gkae1115
-
[2021]
Genome mining methods to discover bioactive natural products.Natural Product Reports38, 11 (2021), 2100–2129. doi:10.1039/d1np00025h
-
[2025]
doi:10.1101/ 2025.06.02.657346
A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products.bioRxiv(2025). doi:10.1101/ 2025.06.02.657346
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.