ArBG replaces flow-based methods with autoregressive models for Boltzmann sampling, showing gains on peptide benchmarks and a 132M-parameter model Robin cutting zero-shot energy error by over 60% on 8-residue systems.
hub
Lin, et al., Evolutionary-scale prediction of atomic-level protein struc- ture with a language model, Science 379 (6637) (2023) 1123–1130
28 Pith papers cite this work, alongside 4,910 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
LOGICA adds context to pretrained biological LMs via logit-space contrastive alignment with gated adapters, improving AUC on held-out drug-resistance mutation ranking from ~0.55 to ~0.65 while preserving token likelihoods.
Introduces first protein sequence dataset for nine Bangladeshi fish species and a deployable hybrid CNN-Transformer model achieving 79.8% accuracy with strong efficiency advantages over ProtBERT.
AMPGAN v3 generates non-canonical AMPs with D-amino acids and modifications using two discriminators for stability, validated with two active candidates in vitro, alongside the PepCraft multi-agent discovery framework.
A policy network learns to choose unmasking order in masked diffusion by reweighting the loss, outperforming random and heuristic baselines on ordering-sensitive tasks.
Activation outliers quantified by γ = ||μ||/||σ|| cause feature death in SAEs via initialization shifts; mean-centering eliminates it.
OOD-GraphLLM is a graphLLM framework that jointly optimizes molecular graph representations and biomedical semantic language representations for out-of-distribution drug synergy prediction.
Screening of 52,000 bioRxiv preprints finds dual-use-adjacent content routinely present in open titles and abstracts, often exceeding risk thresholds.
Dual Triangle Attention achieves effective bidirectional attention with built-in positional inductive bias via dual triangular masks, outperforming standard bidirectional attention on position-sensitive tasks and showing strong masked language modeling results with or without positional embeddings.
AlphaEvolve is an LLM-orchestrated evolutionary coding agent that discovered a 4x4 complex matrix multiplication algorithm using 48 scalar multiplications, the first improvement over Strassen's algorithm in 56 years, plus optimizations for Google data centers and hardware.
A single autoregressive foundation model unifies protein, molecule, and crystal structures into a shared token vocabulary and generates inspectable reasoning traces, achieving SOTA on 67 of 86 scientific tasks.
Frozen embeddings from a pretrained heterogeneous graph transformer over a 6.9M-node metabolic-engineering knowledge graph predict fermentation titers at R²=0.41, outperforming tabular baselines (R²=0.24).
FLaG is a frequency-domain module using FFT, latent queries, and gating that improves token aggregation and shows gains on ESM2 AMP prediction and CIFAR-100 image classification while staying competitive on text tasks.
CaliPPer introduces a distance-based framework that quantifies generalizability, predicts performance metrics like AUROC with low error, and improves predictions on unseen binding data across multiple models and domains.
2D-ProteinRAG is a dual-dimensional RAG framework that incorporates BLAST workflows plus horizontal attribute alignment and vertical homology denoising to improve protein-text QA on both in-distribution and out-of-distribution cases.
A new tree-conditioned edit-flow model for ancestral sequence reconstruction achieves reasonable accuracy on substitution-only evolved sequences and superior localization of changes on natural indel-rich sequences.
Task-aligned supervised geometric stability predicts linear steerability with high accuracy while unsupervised stability detects representational drift earlier and with lower false alarms than CKA or Procrustes.
A dual-encoder TCR-pMHC model with temperature scaling and conformal abstention achieves AUROC 0.813, ECE 0.043, and reduces error from 18.7% to 10.9% at 80% coverage under epitope-held-out evaluation.
RAG-GNN augments GNNs with retrieved literature knowledge via gated fusion to improve functional clustering of 379 proteins in cancer signaling networks, raising silhouette score by 0.093.
Discrete, Gaussian, and simplicial diffusion models for sequences are unified as parameterizations of the Wright-Fisher population genetics model, allowing multi-domain training and stable simplicial diffusion.
PLASMA applies regularized optimal transport with Sinkhorn iterations to produce fast, interpretable residue-level alignments and similarity scores between protein structures.
HSAP introduces a hierarchical framework and sequence-aware algorithm with JIT-optimized NCCL communication to enable correct causal attention computation on hybrid-context packed sequences without limiting parallelism.
Under matched protocols, MAE-3D with channel cross-attention, FFT regularization, and ESM2 alignment outperforms 2D MAE variants on OpenCell protein localization and PPI tasks.
Emyx, a compact flow matching model with EDM reparametrization, outperforms larger protein generators on enzyme design benchmarks with substantially lower training compute.
citing papers explorer
-
Autoregressive Boltzmann Generators
ArBG replaces flow-based methods with autoregressive models for Boltzmann sampling, showing gains on peptide benchmarks and a 132M-parameter model Robin cutting zero-shot energy error by over 60% on 8-residue systems.
-
Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment
LOGICA adds context to pretrained biological LMs via logit-space contrastive alignment with gated adapters, improving AUC on held-out drug-resistance mutation ranking from ~0.55 to ~0.65 while preserving token likelihoods.
-
Protein-Based Fish Species Identification: Dataset, Models, and Insights from Native Bangladeshi Fish
Introduces first protein sequence dataset for nine Bangladeshi fish species and a deployable hybrid CNN-Transformer model achieving 79.8% accuracy with strong efficiency advantages over ProtBERT.
-
Agentic Discovery of Non-Canonical Antimicrobial Peptides with AMPGAN v3
AMPGAN v3 generates non-canonical AMPs with D-amino acids and modifications using two discriminators for stability, validated with two active candidates in vitro, alongside the PepCraft multi-agent discovery framework.
-
Adaptive Order Policies for Masked Diffusion
A policy network learns to choose unmasking order in masked diffusion by reweighting the loss, outperforming random and heuristic baselines on ordering-sensitive tasks.
-
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
Activation outliers quantified by γ = ||μ||/||σ|| cause feature death in SAEs via initialization shifts; mean-centering eliminates it.
-
OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction
OOD-GraphLLM is a graphLLM framework that jointly optimizes molecular graph representations and biomedical semantic language representations for out-of-distribution drug synergy prediction.
-
The Biosecurity Blind Spot: Systematic Dual-use Detection in Open Science Infrastructure
Screening of 52,000 bioRxiv preprints finds dual-use-adjacent content routinely present in open titles and abstracts, often exceeding risk thresholds.
-
Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings
Dual Triangle Attention achieves effective bidirectional attention with built-in positional inductive bias via dual triangular masks, outperforming standard bidirectional attention on position-sensitive tasks and showing strong masked language modeling results with or without positional embeddings.
-
AlphaEvolve: A coding agent for scientific and algorithmic discovery
AlphaEvolve is an LLM-orchestrated evolutionary coding agent that discovered a 4x4 complex matrix multiplication algorithm using 48 scalar multiplications, the first improvement over Strassen's algorithm in 56 years, plus optimizations for Google data centers and hardware.
-
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
A single autoregressive foundation model unifies protein, molecule, and crystal structures into a shared token vocabulary and generates inspectable reasoning traces, achieving SOTA on 67 of 86 scientific tasks.
-
Canopy: A Heterograph Foundation Model for Metabolic Engineering
Frozen embeddings from a pretrained heterogeneous graph transformer over a 6.9M-node metabolic-engineering knowledge graph predict fermentation titers at R²=0.41, outperforming tabular baselines (R²=0.24).
-
Frequency-Domain Latent Attention Gating for Cross-Domain Token Aggregation
FLaG is a frequency-domain module using FFT, latent queries, and gating that improves token aggregation and shows gains on ESM2 AMP prediction and CIFAR-100 image classification while staying competitive on text tasks.
-
CaliPPer: quantifying, predicting and improving AI model performance for binding prediction
CaliPPer introduces a distance-based framework that quantifies generalizability, predicts performance metrics like AUROC with low error, and improves predictions on unseen binding data across multiple models and domains.
-
Unlocking Biological Workflows for Robust Protein-Text Question Answering: A Dual-Dimensional RAG Framework
2D-ProteinRAG is a dual-dimensional RAG framework that incorporates BLAST workflows plus horizontal attribute alignment and vertical homology denoising to improve protein-text QA on both in-distribution and out-of-distribution cases.
-
Tree-Conditioned Edit Flows for Ancestral Sequence Reconstruction
A new tree-conditioned edit-flow model for ancestral sequence reconstruction achieves reasonable accuracy on substitution-only evolved sequences and superior localization of changes on natural indel-rich sequences.
-
The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability
Task-aligned supervised geometric stability predicts linear steerability with high accuracy while unsupervised stability detects representational drift earlier and with lower false alarms than CKA or Procrustes.
-
Calibrated Abstention for Reliable TCR--pMHC Binding Prediction under Epitope Shift
A dual-encoder TCR-pMHC model with temperature scaling and conformal abstention achieves AUROC 0.813, ECE 0.043, and reduces error from 18.7% to 10.9% at 80% coverage under epitope-held-out evaluation.
-
RAG-GNN: Integrating Retrieved Knowledge with Graph Neural Networks for Precision Medicine
RAG-GNN augments GNNs with retrieved literature knowledge via gated fusion to improve functional clustering of 379 proteins in cancer signaling networks, raising silhouette score by 0.093.
-
A Unification of Discrete, Gaussian, and Simplicial Diffusion
Discrete, Gaussian, and simplicial diffusion models for sequences are unified as parameterizations of the Wright-Fisher population genetics model, allowing multi-domain training and stable simplicial diffusion.
-
Fast and Interpretable Protein Substructure Alignment via Optimal Transport
PLASMA applies regularized optimal transport with Sinkhorn iterations to produce fast, interpretable residue-level alignments and similarity scores between protein structures.
-
HSAP: A Hierarchical Sequence-aware Parallelism for Hybrid-Context Generative Models
HSAP introduces a hierarchical framework and sequence-aware algorithm with JIT-optimized NCCL communication to enable correct causal attention computation on hybrid-context packed sequences without limiting parallelism.
-
3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy
Under matched protocols, MAE-3D with channel cross-attention, FFT regularization, and ESM2 alignment outperforms 2D MAE variants on OpenCell protein localization and PPI tasks.
-
Emyx: Fast and efficient all-atom protein generation
Emyx, a compact flow matching model with EDM reparametrization, outperforms larger protein generators on enzyme design benchmarks with substantially lower training compute.
-
Heterogeneous Scientific Foundation Model Collaboration
Eywa enables language-based agentic AI systems to collaborate with specialized scientific foundation models for improved performance on structured data tasks.
-
An Integrated Deep-Learning Framework for Peptide-Protein Interaction Prediction and Target-Conditioned Peptide Generation with ConGA-PepPI and TC-PepGen
An integrated framework with ConGA-PepPI for PepPI prediction and binding-site localization plus TC-PepGen for target-conditioned peptide generation reports 0.839 accuracy and 0.921 AUROC in cross-validation along with 40.39% of generated peptides exceeding native templates on AlphaFold 3 ipTM.
-
Limitations of Sequence-Based Protein Representations for Parkinson's Disease Classification: A Leakage-Free Benchmark
A controlled benchmark shows protein primary sequence representations achieve only moderate discriminative performance (best F1 0.704) for Parkinson's disease classification, with substantial class overlap and no significant differences across methods.
-
Synergistic Benefits of Joint Molecule Generation and Property Prediction
Hyformer jointly models molecule generation and property prediction via alternating attention and joint pre-training, showing synergistic gains in conditional sampling, OOD prediction, and a drug design case for antimicrobial peptides.