REVIEW 5 major objections 6 minor 94 references
Attribution assignment for deep-generative sequence models enables interpretability analysis using positive-only data
T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read GAMA defines sequence-position importance as the difference between integrated gradients of a trained LSTM and a randomly initialized LSTM, and this difference recovers known binding motifs from positive-only antibody data.
desk verdict Promising attribution idea for generative sequence models, but the defining statistic cancels to zero as written and needs to be fixed before the results can be interpreted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the 'IG-tensor': integrated gradients of an LSTM computed for every input-output token pair, structured as a 4D tensor over input encoding, input position, output probability dimension, and output position, then reduced to a 2D per-position profile. GAMA's importance at a token position is the absolute median difference between the trained model's IG profile and a randomly initialized reference model's IG profile, weighted by the sum of the variances of the two IG distributions. The random reference model is the mechanism that isolates learned signal from architecture-specific baseline behavior. IG is computed via 1000 linearly interpolated points between the one-hot input and a zero baseline, following the standard Integrated Gradients procedure, and the whole computation is summarized in Algorithm 1.
What would settle it
Train an LSTM on uniformly random sequences with no implanted motif and run GAMA; if any position is consistently flagged as important across random seeds, the random-model subtraction fails as a null control.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that Integrated Gradients can be turned into an interpretability metric for autoregressive generative models by contrasting a trained model against a randomly initialized reference model. GAMA computes a 4D integrated-gradient tensor over all input tokens, output tokens, and encoding dimensions, reduces it to per-position importance scores, and takes a variance-weighted median absolute difference between the trained and reference models (Equation 1). Benchmarking on 270 synthetic datasets shows that implanted motifs are retrieved with false-negative rates that improve with signal-to-noise ratio and are best for AND-logic motifs, and that retrieval is essentially invariant to motif position. On Absolut!-simulated antibody-antigen data, GAMA attributions correlate with per-position binding affinities (Spearman 0.74 for 5CZV, 0.63 for 4K24, with one-sided p-values below 0.05), and on the experimental Trastuzumab-HER2 dataset GAMA flags positions 103, 104, 105, and 107, recovering three of the four known paratope positions of the unmutated antibody.
Load-bearing premise
The method assumes that subtracting the integrated gradients of a randomly initialized LSTM from those of a trained LSTM removes all architecture-specific and baseline artifacts, leaving only what the model learned from the positive sequences.
Editorial extensions
If this is right
- Generative antibody design pipelines can inspect which positions a model learned to associate with binding, without collecting negative examples.
- GAMA can validate in-silico design strategies by checking that newly generated sequences concentrate importance at known functional positions.
- The method extends in principle to any differentiable autoregressive model, including protein language models and state-space models, making their learned sequence preferences inspectable.
- The operating range is quantified: AND-logic motifs are recovered with false-negative rates below 50% even when the signal fraction is as low as 40%, while XOR-logic motifs require near-clean data.
Reading between the lines
- A natural extension the paper leaves implicit is to use GAMA as a diagnostic on individual sequences; the authors state single-sequence IG noise is too high, so variance reduction across sequences is currently required.
- The same subtraction idea could be tested against explicit null models, such as permuting the training labels or training on shuffled sequences, to check that the random reference is not masking dataset-level biases.
- Because GAMA pools over a dataset, it would likely miss rare or heterogeneous motifs; a position-pair version of the importance metric would be a direct test of whether higher-order dependencies can be recovered.
- If applied to large language models, the reference-model subtraction would need recalibration, since random initialization scales differently with model depth.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces GAMA (Generative Attribution Metric Analysis), a post-hoc attribution method for autoregressive generative sequence models, defined as the difference between Integrated Gradients (IG) computed for a trained LSTM and for a randomly initialized reference LSTM. The method is evaluated on 270 synthetic datasets with implanted motifs under varying noise, motif position, and motif logic; on simulated antibody-antigen binding datasets generated with the Absolut! framework; and on an experimental Trastuzumab-HER2 binder dataset. The authors report strong motif retrieval at high signal-to-noise ratios for AND-logic motifs, position invariance, partial agreement with known binding positions in the Trastuzumab case, and significant correlations with Absolut! binding energies for two of four antigens.
Significance. If the method is correctly specified, GAMA addresses a genuine gap: attribution methods for 1-class generative sequence models, which are relevant for antibody design and other positive-only learning settings. The synthetic benchmark design is a clear strength: 270 controlled conditions with known ground-truth motifs, a Monte Carlo random baseline, and explicit tests of noise, position, and logic. The Absolut! validation is also a thoughtful use of position-wise binding energies as a proxy ground truth, and the paper honestly reports nonsignificant correlations for two of the four antigens. However, the central statistic is not defined precisely enough to be reproducible or even unambiguous, and several claims in the main text conflict with the supplementary material. These issues currently prevent the reader from assessing whether the reported results are produced by the method as described.
major comments (5)
- [Methods, 'Calculation of integrated gradients and structuring into a 4D Tensor'; Equation 1] The reduction of the 4D IG tensor's third axis by 'averaging over all elements' is not a well-defined operation for the reported nonzero results: because the output is a softmax vector, the signed Integrated Gradients summed over output classes are identically zero for every input feature (sum_j ∂p_j/∂x = 0), so a plain average over the output-class axis would produce a zero 2D tensor. The nonzero heatmaps in Figures 4–8 therefore require a different reduction, such as averaging absolute values, averaging logits, or selecting a single output token, and the missing Equation 1 leaves the variance weighting unspecified. The manuscript must state the exact reduction and the full formula for M(s) before the method can be evaluated or reproduced.
- [Methods, 'Synthetic dataset generation'; Results, 'GAMA achieves high motif retrieval performance at high…] The noise parameter is defined inconsistently: the Results define 'signal-to-noise ratio' as the fraction of signal sequences (100% means no noise sequences), while the Methods call the third tuple element 'noise ratio' and define it as the ratio between motif-containing sequences and noise sequences, with values from 1.0 down to 0.1. These definitions are incompatible, because a value of 1.0 would mean equal numbers of signal and noise sequences under the Methods definition, but the Results and Table 2 treat 1.0 as 100% signal. The axes of Figures 5 and 6 and all reported false-negative rates depend on this definition and must be reconciled.
- [Results, 'The performance of GAMA is position-invariant'; Supplementary Figures 3–5] The main-text conclusion that GAMA is position-invariant is contradicted by the supplementary material, whose captions state that 'motif recovery rates are superior when motifs are located at the end of the sequence' (Supplementary Figures 3, 4, and 5). The paper should either reconcile these statements with the main-text claim or soften the position-invariance conclusion to reflect the supplementary evidence.
- [Results, 'GAMA attributions on synthetic antibody binding data correlate with positional contribution to binding…] The claim of significant correlation between GAMA attributions and Absolut! binding energies is supported for only two of four antigens (5CZV p=0.004, 4K24 p=0.018), and no multiple-testing correction is applied; under a Bonferroni threshold of 0.0125, only 5CZV remains significant. The Discussion sentence 'Our results indicate a strong correlation between GAMA and the ground-truth data in simulated antibody binding data' therefore overstates the evidence, and the conclusion should be tempered or the analysis adjusted.
- [Methods, 'GAMA computation algorithm'; Algorithm 1] The choice of a randomly initialized LSTM as the reference and the variance weighting in the similarity function are not justified or ablated. Because 'Importance' is defined as a difference between two IG distributions, the manuscript should show that a single random initialization provides a stable reference and that the variance-weighted absolute median difference outperforms simpler alternatives (for example, raw IG magnitude or IG differences without weighting) on the synthetic benchmarks. As written, the reader cannot determine whether the reported performance is due to the random-reference subtraction, the variance weighting, or the underlying IG signal.
minor comments (6)
- [Results, Figure 5D caption] The caption states 'The false-negative-rate is above 50% for all signal-to-noise conditions ≥80%,' but the main text reports that the false-negative-rate is below 50% for signal-to-noise ratios of 90–100%; this caption should be corrected to match the data.
- [Results and Table 1] The antigen identifier is written inconsistently as '5KN5' in Table 1 and as '5K5N' in the text and Supplementary Figure 7; please unify the spelling.
- [Results, 'GAMA attributions on synthetic antibody binding data'] The typo 'Abolution!' appears twice where 'Absolut!' is intended.
- [Methods, 'Calculation of integrated gradients and structuring into a 4D Tensor'; Supplementary Figure 12] The text says the output probability vector has 21 elements (20 amino acids plus stop token), while Supplementary Figure 12 refers to '22 output dimensions'; the numbers should be aligned.
- [Methods, Equation 1 and Algorithm 1] Equation 1 is referenced but not displayed in the manuscript; the placeholder must be filled with the actual formula for M(s), including the definition of the variance weighting.
- [Full text] No code or data availability statement is provided; for a methods paper, making the GAMA implementation and the synthetic data generation scripts available would substantially aid reproducibility.
Circularity Check
No significant circularity: GAMA's statistic is a baseline-subtraction of integrated gradients with no parameters fitted to the evaluation data, and the paper's self-citations are contextual.
full rationale
GAMA defines Importance as the difference between the integrated gradients of a trained LSTM and a randomly initialized reference LSTM (Figure 2A and Algorithm 1); this is a baseline-subtraction scheme, not a fit. The synthetic benchmarks implant motifs under known ground truth before GAMA is applied, and the Absolut! and Mason comparisons use externally computed binding energies and crystal-structure positions. No GAMA parameter is tuned to any of these evaluation labels, so the reported motif recovery and correlations are not forced by construction. Self-citations occur, for example the Absolut! framework from Robert et al. and the authors' earlier IG-on-antibodies work, but they are contextual or provide an independent simulation dataset; no uniqueness theorem or load-bearing premise is imported from the authors' prior work. The paper explicitly acknowledges limitations, including single-position analysis, aggregated averaging over sequences, high per-sequence IG noise, and an LSTM-only demonstration, which is consistent with empirical, falsifiable claims rather than a derivation that reduces to its inputs. One manuscript defect is that Equation 1 is not shown and the 4D-to-2D reduction is described as averaging over the softmax output dimension; if signed output-class gradients are averaged, the statistic could vanish. This is a correctness and reproducibility issue, not a circularity issue, because it does not make the claimed predictions equal their inputs by construction.
Assumptions & free parameters
assumptions (4)
- domain assumption The difference between integrated gradients of a trained LSTM and a randomly initialized LSTM isolates features learned from the training data, independent of architecture and initialization.
- domain assumption Integrated gradients can be meaningfully computed for a next-token prediction task with a zero-input baseline that is not a valid one-hot encoding.
- domain assumption Position-wise average of integrated gradients over output dimensions and over sequences is an appropriate summary statistic for identifying important positions.
- domain assumption In the Absolut! simulation, amino acid binding energy at a position is a valid proxy for the biological 'importance' of that position for binding.
Cite this review
Pith. "Pith review of Attribution assignment for deep-generative sequence models enables interpretability analysis using positive-only data." pith.science (2026). https://pith.science/paper/NPTFHVWV
@misc{pith2026250623182,
author = {Pith},
title = {Pith review of: Attribution assignment for deep-generative sequence models enables interpretability analysis using positive-only data},
year = {2026},
howpublished = {\url{https://pith.science/paper/NPTFHVWV}},
note = {Machine review of arXiv:2506.23182}
}
read the original abstract
Generative machine learning models offer a powerful framework for therapeutic design by efficiently exploring large spaces of biological sequences enriched for desirable properties. Unlike supervised learning methods, which require both positive and negative labeled data, generative models such as LSTMs can be trained solely on positively labeled sequences, for example, high-affinity antibodies. This is particularly advantageous in biological settings where negative data are scarce, unreliable, or biologically ill-defined. However, the lack of attribution methods for generative models has hindered the ability to extract interpretable biological insights from such models. To address this gap, we developed Generative Attribution Metric Analysis (GAMA), an attribution method for autoregressive generative models based on Integrated Gradients. We assessed GAMA using synthetic datasets with known ground truths to characterize its statistical behavior and validate its ability to recover biologically relevant features. We further demonstrated the utility of GAMA by applying it to experimental antibody-antigen binding data. GAMA enables model interpretability and the validation of generative sequence design strategies without the need for negative training data.
Reference graph
Works this paper leans on
-
[1]
& Schmidhuber, J
Hochreiter, S. & Schmidhuber, J. Long Short-term Memory. Neural computation 9 , 1735–80 (1997)
1997
-
[2]
Gu, A. & Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. Preprint at https://doi.org/10.48550/arXiv.2312.00752 (2024)
-
[3]
Goodfellow, I. J. et al. Generative Adversarial Nets. in Advances in Neural Information Processing Systems (eds. Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N. & Weinberger, K. Q.) vol. 27 (Curran Associates, Inc., 2014)
2014
-
[4]
Kingma, D. P. & Welling, M. Auto-Encoding Variational Bayes. in 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings (eds. Bengio, Y. & LeCun, Y.) (2014)
2014
-
[5]
& Ganguli, S
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N. & Ganguli, S. Deep Unsupervised Learning using Nonequilibrium Thermodynamics. in Proceedings of the 32nd International Conference on Machine Learning 2256–2265 (PMLR, 2015)
2015
-
[6]
Vaswani, A. et al. Attention is All you Need. in Advances in Neural Information Processing Systems vol. 30 (Curran Associates, Inc., 2017)
2017
-
[7]
& Meng-Papaxanthos, L
Kucera, T., Togninalli, M. & Meng-Papaxanthos, L. Conditional generative modeling for de novo protein design with hierarchical functions. Bioinformatics 38 , 3454–3461 (2022)
2022
-
[8]
E., Arnold, F
Wu, Z., Johnston, K. E., Arnold, F. H. & Yang, K. K. Protein sequence design with deep generative models. Current Opinion in Chemical Biology 65 , 18–27 (2021)
2021
Show all 94 references
-
[9]
Chen, Y. et al. Deep generative model for drug design from protein target sequence. Journal of Cheminformatics 15 , 38 (2023)
2023
-
[10]
& Weigt, M
Trinquier, J., Uguzzoni, G., Pagnani, A., Zamponi, F. & Weigt, M. Efficient generative modeling of protein sequences using simple autoregressive models. Nat Commun 12 , 5800 (2021)
2021
-
[11]
Madani, A. et al. Large language models generate functional protein sequences across diverse families. Nat Biotechnol 1–8 (2023) doi:10.1038/s41587-022-01618-2. 25
2023 doi
-
[12]
& Marks, D
Notin, P., Rollins, N., Gal, Y., Sander, C. & Marks, D. Machine learning for functional protein design. Nat Biotechnol 42 , 216–228 (2024)
2024
-
[13]
& Listgarten, J
Hsu, C., Fannjiang, C. & Listgarten, J. Generative models for protein structures and sequences. Nat Biotechnol 42 , 196–199 (2024)
2024
-
[14]
Wang, H. et al. Scientific discovery in the age of artificial intelligence. Nature 620 , 47–60 (2023)
2023
-
[15]
Wu, K. E. et al. Protein structure generation via folding diffusion. Nat Commun 15 , 1059 (2024)
2024
-
[16]
Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620 , 1089–1100 (2023)
2023
-
[17]
Ni, B., Kaplan, D. L. & Buehler, M. J. Generative design of de novo proteins based on secondary-structure constraints using an attention-based diffusion model. Chem 9 , 1828–1849 (2023)
2023
-
[18]
Davidsen, K. et al. Deep generative models for T cell receptor protein sequences. eLife 8 , e46935 (2019)
2019
-
[19]
& Huang, P
Anand, N. & Huang, P. Generative modeling for protein structures
-
[20]
Zrimec, J. et al. Controlling gene expression with deep generative design of regulatory DNA. Nat Commun 13 , 5099 (2022)
2022
- [21]
-
[22]
R., Kim, D
Seo, E., Choi, Y.-N., Shin, Y. R., Kim, D. & Lee, J. W. Design of synthetic promoters for cyanobacteria with generative deep-learning model. Nucleic Acids Research 51 , 7071–7082 (2023)
2023
-
[23]
T., Robson, J
Riley, A. T., Robson, J. M., Ulanova, A. & Green, A. A. Generative and predictive neural networks for the design of functional RNA molecules. Nat Commun 16 , 4155 (2025)
2025
-
[24]
Saka, K. et al. Antibody design using LSTM based deep generative model from phage display library for affinity maturation. Sci Rep 11 , 5852 (2021)
2021
-
[25]
Akbar, R. et al. In silico proof of principle of machine learning-based antibody design at unconstrained scale. mAbs (2022)
2022
-
[26]
Amimeur, T. et al. Designing Feature-Controlled Humanoid Antibody Discovery Libraries Using 26 Generative Adversarial Networks . http://biorxiv.org/lookup/doi/10.1101/2020.04.12.024844 (2020) doi:10.1101/2020.04.12.024844
2020 doi
-
[27]
Shin, J.-E. et al. Protein design and variant prediction using autoregressive generative models. Nat Commun 12 , 2403 (2021)
2021
-
[28]
W., Adler, A
Lim, Y. W., Adler, A. S. & Johnson, D. S. Predicting antibody binders and generating synthetic antibodies using deep learning. mAbs 14 , 2069075 (2022)
2022
- [29]
-
[30]
Hie, B. L. et al. Efficient evolution of human antibodies from general protein language models. Nat Biotechnol 1–9 (2023) doi:10.1038/s41587-023-01763-2
2023 doi
-
[31]
Interpretable Machine Learning
Molnar, C. Interpretable Machine Learning . (Leanpub, 2018)
2018
-
[32]
Liu, Z. et al. Inferring the Effects of Protein Variants on Protein–protein Interactions with Interpretable Transformer Representations. Research research.0219 (2023) doi:10.34133/research.0219
2023 doi
-
[33]
Carter, B., Krog, J., Birnbaum, M. E. & Gifford, D. K. Machine Learning Model Interpretations Explain T Cell Receptor Binding . http://biorxiv.org/lookup/doi/10.1101/2023.08.15.553228 (2023) doi:10.1101/2023.08.15.553228
2023 doi
- [34]
-
[35]
& Yan, Q
Sundararajan, M., Taly, A. & Yan, Q. Axiomatic Attribution for Deep Networks. in Proceedings of the 34th International Conference on Machine Learning 3319–3328 (PMLR, 2017)
2017
-
[36]
A., Sulam, J
Ruffolo, J. A., Sulam, J. & Gray, J. J. Antibody structure prediction using interpretable deep learning. Patterns 3 , 100406 (2022)
2022
-
[37]
Senior, A. W. et al. Improved protein structure prediction using potentials from deep learning. Nature 577 , 706–710 (2020). 27
2020
-
[38]
& Unterthiner, T
Preuer, K., Klambauer, G., Rippmann, F., Hochreiter, S. & Unterthiner, T. Interpretable Deep Learning in Drug Discovery. in Explainable AI: Interpreting, Explaining and Visualizing Deep Learning (eds. Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K. & Müller, K.-R.) 331–345...
2019 doi
-
[39]
& Okuno, Y
Ishida, S., Terayama, K., Kojima, R., Takasu, K. & Okuno, Y. Prediction and Interpretable Visualization of Retrosynthetic Reactions Using Graph Convolutional Networks. J. Chem. Inf. Model. 59 , 5026–5033 (2019)
2019
-
[40]
S., Sampson, A
Nearing, G. S., Sampson, A. K., Kratzert, F. & Frame, J. Post-Processing a Conceptual Rainfall-Runoff Model with an LSTM. (2020)
2020
-
[41]
A., Ehsani, M
De la Fuente, L. A., Ehsani, M. R., Gupta, H. V. & Condon, L. E. Towards Interpretable LSTM-based Modelling of Hydrological Systems. EGUsphere 1–29 (2023) doi:10.5194/egusphere-2023-666
2023 doi
- [42]
-
[43]
& Shen, H.-B
Lin, Y., Pan, X. & Shen, H.-B. lncLocator 2.0: a cell-line-specific subcellular localization predictor for long non-coding RNAs with interpretable deep learning. Bioinformatics 37 , 2308–2316 (2021)
2021
-
[44]
Robert, P. A. et al. Unconstrained generation of synthetic antibody–antigen structures to guide machine learning methodology for antibody specificity prediction. Nat Comput Sci 2 , 845–865 (2022)
2022
-
[45]
Ursu, E. et al. Training data composition determines machine learning generalization and biological rule discovery. Preprint at https://doi.org/10.1101/2024.06.17.599333 (2024)
2024 doi
-
[46]
Widrich, M. et al. Modern hopfield networks and attention for immune repertoire classification. in Proceedings of the 34th International Conference on Neural Information Processing Systems 18832–18845 (Curran Associates Inc., Red Hook, NY, USA, 2020)
2020
-
[47]
Arras, L. et al. Explaining and Interpreting LSTMs. in Explainable AI: Interpreting, Explaining and Visualizing Deep Learning (eds. Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K. & Müller, K.-R.) 211–238 (Springer International Publishing, Cham, 2019). 28 doi:10.1007/978-...
2019 doi
- [48]
-
[49]
Sequential Integrated Gradients: a simple but effective method for explaining language models
Enguehard, J. Sequential Integrated Gradients: a simple but effective method for explaining language models. in Findings of the Association for Computational Linguistics: ACL 2023 (eds. Rogers, A., Boyd-Graber, J. & Okazaki, N.) 7555–7565 (Association for Computational Linguis...
2023 doi
-
[50]
& Bosnić, Z
Papič, A., Kononenko, I. & Bosnić, Z. Conditional generative positive and unlabeled learning. Expert Systems with Applications 224 , 120046 (2023)
2023
-
[51]
& Jha, S
Pramanik, V., Maliha, M. & Jha, S. K. Enhancing Integrated Gradients Using Emphasis Factors and Attention for Effective Explainability of Large Language Models. (2024)
2024
- [52]
-
[53]
Zhou, W., Adel, H., Schuff, H. & Vu, N. T. Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings. in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLI...
2024
- [54]
-
[55]
& Hess, M
Treppner, M., Binder, H. & Hess, M. Interpretable generative deep learning: an illustration with single cell gene expression data. Hum Genet 141 , 1481–1498 (2022)
2022
-
[56]
A., Adebayo, J., Bravo, H
Ismail, A. A., Adebayo, J., Bravo, H. C., Ra, S. & Cho, K. Concept Bottleneck Generative Models
-
[57]
Liu, W. et al. Towards Visually Explaining Variational Autoencoders. in 8639–8648 (IEEE Computer Society, 2020). doi:10.1109/CVPR42600.2020.00867
2020
-
[58]
Z., Glassman, E
Ross, A., Chen, N., Hang, E. Z., Glassman, E. L. & Doshi-Velez, F. Evaluating the Interpretability of Generative Models by Interactive Reconstruction. in Proceedings of the 2021 CHI Conference on 29 Human Factors in Computing Systems 1–15 (Association for Computing Machinery, ...
2021
-
[59]
& Hotho, A
Tritscher, J., Krause, A. & Hotho, A. Feature relevance XAI in anomaly detection: Reviewing approaches and challenges. Front. Artif. Intell. 6 , (2023)
2023
-
[60]
Sidorczuk, K. et al. Benchmarks in antimicrobial peptide prediction are biased due to the selection of negative data. Briefings in Bioinformatics 23 , bbac343 (2022)
2022
-
[61]
Montemurro, A., Jessen, L. E. & Nielsen, M. NetTCR-2.1: Lessons and guidance on how to develop models for TCR specificity predictions. Frontiers in Immunology 13 , (2022)
2022
-
[62]
& White, A
Ansari, M. & White, A. D. Learning peptide properties with positive examples only. Digital Discovery 3 , 977–986 (2024)
2024
-
[63]
Akbar, R. et al. A compact vocabulary of paratope-epitope interactions enables predictability of antibody-antigen binding. Cell Reports 34 , 108856 (2021)
2021
-
[64]
Narayanan, H. et al. Machine Learning for Biologics: Opportunities for Protein Engineering, Developability, and Formulation. Trends in Pharmacological Sciences 42 , 151–165 (2021)
2021
-
[65]
& Rayalu, G
Ramanjineyulu, B., Pagadala, B. & Rayalu, G. M. Statistical Inference in Autoregressive Models: Estimation of Autoregressive Models . (LAP LAMBERT Academic Publishing, S.l., 2013)
2013
-
[66]
Lamb, A. M. et al. Professor Forcing: A New Algorithm for Training Recurrent Networks. in Advances in Neural Information Processing Systems (eds. Lee, D., Sugiyama, M., Luxburg, U., Guyon, I. & Garnett, R.) vol. 29 (Curran Associates, Inc., 2016)
2016
-
[67]
H., Greiff, V., Karatt-Vellatt, A., Muyldermans, S
Laustsen, A. H., Greiff, V., Karatt-Vellatt, A., Muyldermans, S. & Jenkins, T. P. Animal Immunization, in Vitro Display Technologies, and Machine Learning for Antibody Discovery. Trends Biotechnol 39 , 1263–1273 (2021)
2021
-
[68]
Drug development: the journey of a medicine from lab to shelf
Torjesen, I. Drug development: the journey of a medicine from lab to shelf. The Pharmaceutical Journal https://pharmaceutical-journal.com/article/feature/drug-development-the-journey-of-a-medicine-from -lab-to-shelf (2015). 30
2015
-
[69]
Vu, M. H. et al. Linguistics-based formalization of the antibody language as a basis for antibody language models. Nat Comput Sci 4 , 412–422 (2024)
2024
-
[70]
Karim, M. R. et al. Explainable AI for Bioinformatics: Methods, Tools and Applications. Briefings in Bioinformatics 24 , bbad236 (2023)
2023
-
[71]
Scheffer, L. et al. Predictability of antigen binding based on short motifs in the antibody CDRH3. Brief Bioinform 25 , bbae537 (2024)
2024
-
[72]
Mason, D. M. et al. Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning. Nat Biomed Eng 5 , 600–612 (2021)
2021
-
[73]
Navin, N. E. Cancer genomics: one cell at a time. Genome Biology 15 , 452 (2014)
2014
-
[74]
& Schmidhuber, J
Bengio, Y., Frasconi, P. & Schmidhuber, J. Gradient Flow in Recurrent Nets: the Difficulty of Learning Long-Term Dependencies. A Field Guide to Dynamical Recurrent Neural Networks (2003)
2003
-
[75]
& Davis, J
Bekker, J. & Davis, J. Learning from positive and unlabeled data: a survey. Mach Learn 109 , 719–760 (2020)
2020
-
[76]
& Lee, W
Zhang, D. & Lee, W. S. Learning classifiers without negative examples: A reduction approach. in 2008 Third International Conference on Digital Information Management 638–643 (IEEE, London, United Kingdom, 2008). doi:10.1109/ICDIM.2008.4746761
2008
- [77]
-
[78]
Tang, X. et al. A Survey of Generative AI for de novo Drug Design: New Frontiers in Molecule and Protein Generation. Preprint at http://arxiv.org/abs/2402.08703 (2024)
2024 arXiv
-
[79]
Explainable Generative AI (GenXAI): A Survey, Conceptualization, and Research Agenda
Schneider, J. Explainable Generative AI (GenXAI): A Survey, Conceptualization, and Research Agenda. Preprint at http://arxiv.org/abs/2404.09554 (2024)
2024 arXiv
-
[80]
S., Farmery, J
Leem, J., Mitchell, L. S., Farmery, J. H. R., Barton, J. & Galson, J. D. Deciphering the language of antibodies using self-supervised learning. PATTER 3 , (2022)
2022
- [81]
-
[82]
Beck, M. et al. xLSTM: Extended Long Short-Term Memory. Preprint at http://arxiv.org/abs/2405.04517 (2024)
2024 arXiv
-
[83]
Wilman, W. et al. Machine-designed biotherapeutics: opportunities, feasibility and advantages of deep learning in computational antibody discovery. Briefings in Bioinformatics 23 , bbac267 (2022)
2022
-
[84]
Akbar, R. et al. Progress and challenges for the machine learning-based design of fit-for-purpose monoclonal antibodies. mAbs 14 , 2008790 (2022)
2022
-
[85]
M., Kinney, J
Adams, R. M., Kinney, J. B., Walczak, A. M. & Mora, T. Epistasis in a Fitness Landscape Defined by Antibody-Antigen Binding Free Energy. Cell Syst 8 , 86-93.e3 (2019)
2019
-
[86]
Schober, K., Buchholz, V. R. & Busch, D. H. TCR repertoire evolution during maintenance of CMV-specific T-cell populations. Immunological Reviews 283 , 113–128 (2018)
2018
-
[87]
Machine Learning Analysis of Naïve B-Cell Receptor Repertoires Stratifies Celiac Disease Patients and Controls
Yaari, G. Machine Learning Analysis of Naïve B-Cell Receptor Repertoires Stratifies Celiac Disease Patients and Controls. Frontiers in Immunology 12 , (2021)
2021
-
[88]
Kingma, D. P. & Ba, J. Adam: A Method for Stochastic Optimization. in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings (eds. Bengio, Y. & LeCun, Y.) (2015)
2015
-
[89]
& Kinney, J
Tareen, A. & Kinney, J. B. Logomaker: Beautiful sequence logos in python. 635029 Preprint at https://doi.org/10.1101/635029 (2019)
2019 doi
-
[90]
Waskom, M. L. seaborn: statistical data visualization. Journal of Open Source Software 6 , 3021 (2021)
2021
- [91]
-
[92]
Harris, C. R. et al. Array programming with NumPy. Nature 585 , 357–362 (2020)
2020
-
[93]
Data Structures for Statistical Computing in Python
McKinney, W. Data Structures for Statistical Computing in Python. Proceedings of the 9th Python in Science Conference 56–61 (2010) doi:10.25080/Majora-92bf1922-00a
2010 doi
-
[94]
Virtanen, P. et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods 17 , 261–272 (2020). 32 Supplementary Material Supplementary Figure 1 A learning rate of 0.00001 is effective for training the LSTM on the synthetic dataset. Evaluating the in...
2020
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.