{"id":"17920a53-341c-4215-bec1-27bc8f73ffd0","arxiv_id":"2507.07666","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive review of machine learning approaches for enzyme mining, ending with a proposed but untested modular pipeline.","lead":"This paper surveys how machine learning tools predict enzyme functions and properties, and sketches a proposed AI-guided pipeline for enzyme discovery. It is a useful map of a fast-moving field for researchers and companies working with biocatalysts.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'establish scalable and predictive framework' is undercut by the paper's own Section 3.3 concession that models underperform on de novo/low-homology sequences, and the proposed Section 4.1 pipeline has no end-to-end validation.","rationale":"The reader's weakest assumption is essentially the same load-bearing issue: the pipeline's accuracy on uncharacterized, low-homology sequences. I partially agree, but I would sharpen it from an open assumption into an internal tension: the abstract claims the framework is established while Section 3.3 concedes that generalization to de novo sequences remains a common failure across predictors. That internal inconsistency is the strongest concern. The paper is explicitly a review with no new datasets or models, and the reader's UNVERDICTED verdict appropriately reflects that there is no original research claim to verify. The case studies with experimental validation do provide partial empirical support for ML-guided mining, so this is not a claim that the approach is worthless; the problem is the global 'establish' claim. Because the review remains useful as a synthesis, and because the reader already judged the contribution as unverifiable rather than as a failed research claim, my recommendation is UNCHANGED. The abstract should be tempered in revision, but the verdict does not need to move.","tokens_in":36480,"tokens_out":3842,"duration_ms":44592,"concrete_test":"Apply the Section 4.1 pipeline end-to-end to a set of enzymes characterized after the training cutoffs of the cited models (e.g., newly deposited BRENDA or recent literature entries) that share less than 30% sequence identity with any training sequence. Compute top-k precision of the multi-objective ranking against experimentally confirmed activity, and compare it with a simple homology/SSN baseline on the same pool. If the ML pipeline's top-k hit rate is not significantly higher than the baseline, the abstract's claim that the framework is 'established' as scalable and predictive fails; if it is higher, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the abstract's assertion that 'these developments establish ML-guided enzyme mining as a scalable and predictive framework.' For that to be true, the component models must prioritize uncharacterized, low-homology metagenomic sequences that survive experimental validation. The paper's own Section 3.3 states the opposite condition: 'models trained solely on annotated sequences often underperform on de novo sequences without close homologs, a common scenario across all predictors so far.' Section 4.2 repeats that dataset bias leads to 'reduced generalizability to novel or low-homology sequences.' The proposed Section 4.1 pipeline is a conceptual diagram; the Data Availability statement says the work produces no datasets, training models, or novel results, and no end-to-end run of the pipeline is reported. The cited case studies (e.g., PU-EPP, DeepMineLys) are positive examples, but they do not establish that the assembled pipeline outperforms homology-based screening across diverse enzyme families. Thus the strongest reading of the abstract is not supported by the presented evidence, and the paper's own limitations sections point to the unresolved condition. This is an internal overclaim, not an external disagreement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a review of machine-learning approaches for enzyme mining, covering functional annotation (EC numbers, GO terms, substrate specificity) and prediction of enzymatic properties (kinetic parameters, thermophilicity, optimal temperature/pH, solubility). It surveys a large number of tools and models, supplemented by eight tables in the Supplementary Material, and presents several case studies of ML-guided enzyme discovery, including PET hydrolases, mycotoxin-degrading enzymes, terpene synthases, and phage lysins. The paper also discusses current limitations—data bias, generalizability, interpretability—and proposes a modular, ML-guided enzyme-mining pipeline in Section 4.1. The abstract concludes that these developments 'establish ML-guided enzyme mining as a scalable and predictive framework for uncovering novel biocatalysts.'","tokens_in":36694,"tokens_out":5097,"duration_ms":53897,"significance":"The main value of the paper is as a structured survey: it organizes a rapidly growing literature, gives a broad inventory of models and their input/output types, and the supplementary tables provide a useful reference for practitioners. The authors are also explicit about several persistent challenges, which is a strength for a review in this area. However, the paper's headline conclusion overstates what the evidence supports. The proposed pipeline in Section 4.1 is not implemented or evaluated, and the paper's own Section 3.3 acknowledges that current predictors underperform on de novo and low-homology sequences. If the conclusion is tempered to present ML-guided enzyme mining as a promising and accelerating direction rather than an established scalable framework, and if the Section 4.1 pipeline is clearly labeled as a proposal with a discussion of validation needs, the manuscript would be a solid and useful contribution.","major_comments":[{"comment":"The abstract's assertion that 'these developments establish ML-guided enzyme mining as a scalable and predictive framework' is not supported by the body of the paper. Section 3.3 states that models trained solely on annotated sequences 'often underperform on de novo sequences without close homologs, a common scenario across all predictors so far,' and Section 4.2 repeats that dataset bias leads to 'reduced generalizability to novel or low-homology sequences.' The case studies in Section 3.4 are selected successes, not a systematic demonstration of scalability or predictive reliability. Please revise the abstract and the concluding statements in Sections 4.2 and 5 to frame ML-guided enzyme mining as a promising direction or an emerging capability, and make the claimed status of the framework explicit rather than asserted.","section":"Abstract; §3.3; §4.2"},{"comment":"The proposed ML-guided pipeline is presented as an integrated modular framework, but it is not implemented or evaluated end to end, and the Data Availability statement confirms that the work 'does not produce, collect or use datasets, training models, and novel results.' This is acceptable for a perspective paper, but the manuscript should clearly state in Section 4.1 that the pipeline is a conceptual proposal. As written, the benefits of some steps—for example, latent-space expansion and multi-objective scoring—are described in confident terms as if their effectiveness were established. Please add a discussion of the validation steps needed to test this pipeline, such as benchmark datasets, comparison against homology-based ranking baselines, and criteria for evaluating closed-loop improvement.","section":"§4.1; Data Availability"},{"comment":"Several performance claims in the main text are reported without sufficient evaluation context, which weakens the paper's reliability as a survey. For instance, Section 3.2.2 states that ThermoFinder 'pushed predictive accuracy beyond 98%' without specifying the dataset, the split, or the comparison baseline, and the DeepMineLys case study reports an F1-score on an independent validation set but does not describe how that set was constructed or whether homology-based screening was compared. Given the paper's central claim of a 'predictive framework,' the review should either provide basic evaluation context for the highlighted tools or explicitly caution that reported accuracies are not directly comparable across different benchmarks and datasets.","section":"§3.2.2; §3.4"}],"minor_comments":[{"comment":"The sentence 'Feedback from experimental results is can be used to refine selection criteria' contains a typo; it should read 'Feedback from experimental results can be used'.","section":"§2"},{"comment":"The phrase 'Seq2Topt extents applicability' should be 'Seq2Topt extends applicability', and 'adaption' should be 'adaptation'.","section":"§3.2.2"},{"comment":"There are several typographical issues in this section: 'over representation' should be 'overrepresentation', 'such asKm andkcat' should be 'such as Km and kcat', and 'underperform onde novo sequences' should be 'underperform on de novo sequences'.","section":"§3.3"},{"comment":"The sentence beginning 'These imbalances biases toward learning algorithms' is ungrammatical; consider rewriting as 'These imbalances bias learning algorithms toward dominant enzyme families and reduce model performance on understudied or novel proteins.'","section":"§3.3"},{"comment":"There is a missing space in 'applications(Ariaeenejad et al., 2024)'; it should read 'applications (Ariaeenejad et al., 2024)'.","section":"§3.2.1"},{"comment":"The phrase 'undocumented substrate ranges further compromise' is missing a comma; it should read 'undocumented substrate ranges, further compromise'.","section":"§3.3"},{"comment":"The phrase 'the limitations of available training data' could be made more specific; based on the surrounding text, 'the composition and coverage of available training data' would be more accurate.","section":"§4.2"},{"comment":"A summary table of the case studies in Section 3.4—including target enzyme family, ML method, training set size, validation metric, and number of experimentally confirmed hits—would improve the readability and comparative value of the review.","section":"§3.4"}],"recommendation":"major_revision","confidential_remarks":"This is a competent review with a useful tool inventory, but the abstract clearly overstates the evidence, and the main proposal is unvalidated. The authors should be asked to align the title/abstract/conclusion with the actual scope. I also note that a number of cited items are from the authors' own groups, including a bioRxiv preprint used as a key case study (Medina-Ortiz et al. 2025); this is not inherently problematic for a review, but the dependence on unpublished work should be checked and possibly disclosed. The manuscript fits the scope of a review/perspective in q-bio.BM."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this is a review, not a research paper, and the authors say so in the Data Availability statement: no datasets, training models, or novel results. What it does well is organize the current ML-for-enzyme-mining landscape—EC/GO prediction, substrate specificity, kinetic parameters, thermostability, pH/temperature optima, solubility—with a useful set of supplementary tables. That is a genuinely useful map for someone entering the field. The case studies (PU-EPP, DeepMineLys, terpene synthases) are well chosen and do show that ML-guided mining can produce validated hits.\n\nThe soft spot is the abstract's closing claim that 'these developments establish ML-guided enzyme mining as a scalable and predictive framework.' That is not established by the paper's own evidence. Section 3.3 is explicit that models 'trained solely on annotated sequences often underperform on de novo sequences without close homologs, a common scenario across all predictors so far,' and Section 4.2 repeats that dataset bias 'reduces generalizability to novel or low-homology sequences.' The proposed pipeline in Section 4.1 is a conceptual diagram assembled from existing tools; no end-to-end run is reported. So the strongest reading of the abstract is an internal overclaim—the body is more careful than the abstract. That is a real flaw, but it's confined to the framing; it doesn't invalidate the review's descriptive content.\n\nI'd also note that the authors cite their own earlier bioRxiv preprint in the PET-hydrolase case study, but that's not a problem in a review; the other case studies come from independent groups.\n\nWho is this for? A graduate student or researcher wanting a quick, reasonably current panorama of ML predictors in biocatalysis. It is not for someone seeking a new method or benchmark. As a review it deserves a serious referee—an editor should send it out rather than desk-reject—because the coverage is broad and the supplementary material is carefully assembled. The referee should ask the authors to temper the abstract and make the Section 4.1 proposal explicitly a proposal, not an accomplished framework. With that revision, it's a sound, citable review.","headline":"A serviceable, wide-ranging review of ML tools for enzyme mining that is honestly self-limited in the body but overstates itself in the abstract.","tokens_in":37224,"tokens_out":1611,"would_cite":true,"duration_ms":17756,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine learning turns enzyme mining into a scalable, predictive discovery framework.","keywords":["enzyme mining","machine learning","functional annotation","protein language models","biocatalyst discovery","metagenomics","enzyme property prediction","EC number prediction"],"falsifier":"Take a large metagenomic enzyme pool, run the proposed pipeline for a specific target reaction, and experimentally test the top-ranked candidates while also testing candidates chosen by sequence similarity and by random selection. If the ML-ranked set does not show a clearly higher experimental hit rate across several enzyme families, the claim that ML makes enzyme mining scalable and predictive is not supported.","tokens_in":36290,"feed_emoji":"🧬","tokens_out":6944,"duration_ms":69519,"temperature":0.7,"pith_summary":"Enzyme mining is the effort to find useful biocatalysts inside the enormous, mostly unannotated space of protein sequences. This review argues that machine learning has turned that effort into a scalable and predictive framework: models can now assign Enzyme Commission numbers, Gene Ontology terms, substrate specificity, kinetic parameters, optimal temperature and pH, and solubility directly from sequence. The authors survey the model landscape and selected case studies, and propose a modular pipeline in which protein language model embeddings, functional classifiers, property estimators, and multi-objective scoring prioritize candidates for experimental validation. They also list the obstacles that remain, above all data scarcity, dataset bias, and poor generalization to sequences without close homologs. The payoff, if the claim holds, is that computational prioritization can replace much of the slow cultivation-based screening in biocatalyst discovery.","feed_headline":"ML can mine novel enzymes from uncharacterized sequence space","feed_subtitle":"Review shows ML models for enzyme function and properties, and a pipeline to validate candidates at scale.","key_machinery":"The load-bearing mechanism is the protein language model embedding: a neural network trained on vast numbers of protein sequences converts each enzyme into a vector that encodes sequence-function relationships, and these vectors define a latent space in which clustering reveals underexplored or functionally divergent regions. Around that core, the paper assembles functional classifiers (EC number, GO term, substrate specificity) and property estimators (kinetic constants, optimal temperature and pH, solubility) whose outputs feed a multi-objective scoring function that ranks candidates by novelty, predicted promiscuity, and user-defined application traits. That combined machinery lets the pipeline prioritize sequences that homology search would miss, and the closed loop of experimental feedback is what the authors claim will make discovery autonomous.","core_discovery":"The paper's central claim is that ML-guided enzyme mining now offers a scalable and predictive route to novel biocatalysts, in contrast to cultivation-based discovery and homology-only annotation. To support this, it catalogues machine learning models across functional annotation (EC numbers, GO terms, substrate specificity) and property prediction (Km and kcat, thermostability and thermophilicity, optimal temperature and pH, solubility), and shows through case studies of PET hydrolases, mycotoxin-degrading enzymes, terpene synthases, and phage lysins that predicted candidates repeatedly survive experimental validation. The proposed framework makes the claim concrete: build a targeted enzyme pool, embed sequences with pretrained protein language models, analyze the latent space for underexplored clusters, annotate with classifiers and estimators, rank by novelty, promiscuity, and application traits, validate experimentally, and feed results back into the models. The authors are careful to frame this as the direction of the field rather than a finished system; the unresolved limits they name are data scarcity, dataset bias, limited interpretability, and weak generalization to de novo or low-homology sequences.","pith_inferences":["If this claim is right, the field's critical bottleneck shifts from discovery to benchmarking: standardized metagenomic pools with measured experimental hit rates per enzyme family would be the infrastructure needed to decide which models truly generalize.","The same latent-space embeddings used to mine enzymes could double as fitness landscapes for directed evolution or generative design, linking discovery to engineering more tightly than the paper's pipeline states.","One testable extension is to apply the pipeline to enzymes whose substrates are rare or poorly annotated, where negative data are scarce; success there would strengthen the case more than another benchmark on well-studied families.","Because most property predictors are trained on a small number of expression hosts, transferring the framework to non-model expression systems is an open risk the paper flags only implicitly."],"forward_implications":["If ML-guided prioritization works at scale, experimental screening effort shifts from searching broadly to validating a shortlist, cutting cost and time per discovery.","Sequence-only models make enzymes from uncultured and extremophilic microbes accessible to mining, since metagenomic DNA can be analyzed without cultivation.","Including substrate and product information in models should improve prediction of promiscuous enzymes and complete reaction outcomes, not just single enzyme-substrate pairs.","Multi-task and multi-modal learning, plus positive-unlabeled methods, offer concrete routes around the data-scarcity and annotation-bias problems the paper documents.","A closed-loop pipeline that returns experimental results into model training should progressively improve generalizability and move enzyme discovery toward autonomous platforms."],"supporting_citations":[{"why":"CLEAN uses contrastive learning to predict EC numbers, the functional annotation backbone of the proposed pipeline.","marker":"Yu et al., 2023b"},{"why":"ESP provides the general enzyme-substrate specificity prediction used to rank candidate substrates.","marker":"Kroll et al., 2023"},{"why":"CatPred predicts kinetic parameters with structural features, supporting property-based prioritization.","marker":"Boorla and Maranas, 2025"},{"why":"DeepMineLys is the phage-lysin case study showing ML candidates validated experimentally.","marker":"Fu et al., 2024"},{"why":"ESM embeddings are the pretrained protein language model representations that define the latent space.","marker":"Rives et al., 2021"},{"why":"Seq2Topt predicts enzyme optimal temperature from sequence, a property estimator used in the pipeline.","marker":"Qiu et al., 2025"},{"why":"PU-EPP is the mycotoxin-degradation case study where most top-ranked enzymes were experimentally confirmed.","marker":"Zhang et al., 2023"},{"why":"EpHod predicts optimal pH with interpretable residue-level attention, an example of explainable property prediction.","marker":"Gado et al., 2025"}],"fun_headline_variants":["AI mines novel enzymes from unexplored sequence space","Machine learning unearths biocatalysts from dark proteomes","Enzyme discovery: ML finds candidates that pass validation","ML-guided enzyme mining: from sequences to working catalysts","AI accelerates enzyme mining, but data gaps remain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that machine-learning predictions on uncharacterized sequences that look unlike any known enzyme are accurate enough that the top-ranked candidates really do work when tested in the lab.","fun_headline_variants_meta":{"raw":{"variants":["AI mines novel enzymes from unexplored sequence space","Machine learning unearths biocatalysts from dark proteomes","Enzyme discovery: ML finds candidates that pass validation","ML-guided enzyme mining: from sequences to working catalysts","AI accelerates enzyme mining, but data gaps remain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00083,"raw_usage":{"total_tokens":3615,"prompt_tokens":927,"completion_tokens":2688,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":2612}},"tokens_in":543,"tokens_out":2688,"duration_ms":24067,"temperature":1.0,"reasoning_tokens":2612,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:34:32.343485+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a large metagenomic enzyme pool, run the proposed pipeline for a specific target reaction, and experimentally test the top-ranked candidates while also testing candidates chosen by sequence similarity and by random selection. If the ML-ranked set does not show a clearly higher experimental hit rate across several enzyme families, the claim that ML makes enzyme mining scalable and predictive is not supported.","supporting_citations":[],"review_version":1}