A-CODE presents a fully atomic one-stage multimodal diffusion model for protein co-design that claims superior unconditional generation performance over prior one- and two-stage models plus a tenfold success-rate gain on hard binder-design tasks.
hub Mixed citations
Evolutionary-scale prediction of atomic-level protein structure with a language model.Science, 379(6637):1123–1130
Mixed citation behavior. Most common role is method (60%).
hub tools
citation-role summary
citation-polarity summary
years
2026 17representative citing papers
CDM amortizes SMC inference for reward-tilted discrete diffusion by training a parameterized twist function on contrastive samples with closed-form kernels.
BioXArena benchmarks LLM agents on generating end-to-end ML pipelines for 76 multi-modal biomedical tasks, with MLEvolve plus Gemini-3.1-Pro scoring highest at 0.666.
A genome-conditioned 4B LLM agent predicts microbial life boundaries and matches larger frontier models via token fusion, tool use, and a counterfactual gene-grounding reward.
Germline-absorbing discrete diffusion uses the germline sequence as the absorbing state to reduce germline bias in antibody modeling, raising non-germline residue prediction accuracy from 26% to 46% and improving conditional generation tradeoffs over EvoProtGrad.
A five-level physics-informed hierarchical GNN with bidirectional cross-scale fusion substantially improves hard fold classification and reaction-class prediction over geometric and sequence baselines.
Post-training stages reshape generalization in biological reasoning models distinctly: CPT aligns with biological language, SFT boosts ID performance but causes OOD to peak early and decline, while RL on strong SFT checkpoints can recover OOD generalization.
VQ-Atom discretizes local atomic environments into semantic tokens via vector quantization, reaching AUROC 0.79 on KIBA drug-target interaction prediction while enabling 3x faster downstream training than continuous representations.
Yeti is a compact tokenizer for protein structures that delivers strong codebook use, token diversity, and reconstruction while enabling from-scratch multimodal generation of plausible sequences and structures with 10x fewer parameters than ESM3.
SGRPO is a GRPO-style framework that constructs set-level diversity rewards via supergroup sampling and leave-one-out redistribution to expand the utility-diversity Pareto frontier in biomolecular design tasks.
DPLM-Evo introduces an evolutionary discrete diffusion framework with explicit edit prediction and contextual noising that claims SOTA single-sequence mutation effect prediction on ProteinGym while supporting variable-length evolution simulation.
MKGR integrates region-aware sequence encoding with four protein-centered KGs via graph attention, bridge reconstruction, and pair gating to outperform baselines in cold-start PPI prediction.
TCR-SRIM uses structure regularization and contact prototypes for interpretable TCR-epitope binding prediction, reports SOTA performance on TCR-XAI, and finds generated structures produce less accurate interaction patterns than experimental ones.
A new benchmarking framework shows virtual cell models overestimate performance on standard tests, drop sharply on unseen contexts and perturbations, and produce inconsistent rankings across metrics.
Polyformer generates sequence- and temperature-dependent conformational ensembles for proteins that agree with molecular dynamics simulations.
MolClaw deploys a hierarchical skill architecture to reach state-of-the-art results on a new benchmark of multi-step drug discovery tasks.
SubQuad reports near-subquadratic immune-repertoire analysis with fairness-aware clustering, but its central performance and coverage claims are not supported by reproducible artifacts.
citing papers explorer
-
A-CODE: Fully Atomic Protein Co-Design with Unified Multimodal Diffusion
A-CODE presents a fully atomic one-stage multimodal diffusion model for protein co-design that claims superior unconditional generation performance over prior one- and two-stage models plus a tenfold success-rate gain on hard binder-design tasks.
-
Contrastive Distribution Matching for Amortized Sequential Monte Carlo in Discrete Diffusion
CDM amortizes SMC inference for reward-tilted discrete diffusion by training a parameterized twist function on contrastive samples with closed-form kernels.
-
BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks
BioXArena benchmarks LLM agents on generating end-to-end ML pipelines for 76 multi-modal biomedical tasks, with MLEvolve plus Gemini-3.1-Pro scoring highest at 0.666.
-
GGBound: A Genome-Grounded Agent for Microbial Life-Boundary Prediction
A genome-conditioned 4B LLM agent predicts microbial life boundaries and matches larger frontier models via token fusion, tool use, and a counterfactual gene-grounding reward.
-
Conditional generation of antibody sequences with classifier-guided germline-absorbing discrete diffusion
Germline-absorbing discrete diffusion uses the germline sequence as the absorbing state to reduce germline bias in antibody modeling, raising non-germline residue prediction accuracy from 26% to 46% and improving conditional generation tradeoffs over EvoProtGrad.
-
PRIME: Protein Representation via Physics-Informed Multiscale Equivariant Hierarchies
A five-level physics-informed hierarchical GNN with bidirectional cross-scale fusion substantially improves hard fold classification and reaction-class prediction over geometric and sequence baselines.
-
How Post-Training Shapes Biological Reasoning Models
Post-training stages reshape generalization in biological reasoning models distinctly: CPT aligns with biological language, SFT boosts ID performance but causes OOD to peak early and decline, while RL on strong SFT checkpoints can recover OOD generalization.
-
VQ-Atom: Semantic Discretization of Local Atomic Environments for Molecular Representation Learning
VQ-Atom discretizes local atomic environments into semantic tokens via vector quantization, reaching AUROC 0.79 on KIBA drug-target interaction prediction while enabling 3x faster downstream training than continuous representations.
-
Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation
Yeti is a compact tokenizer for protein structures that delivers strong codebook use, token diversity, and reconstruction while enabling from-scratch multimodal generation of plausible sequences and structures with 10x fewer parameters than ESM3.
-
Pushing Biomolecular Utility-Diversity Frontiers with Supergroup Relative Policy Optimization
SGRPO is a GRPO-style framework that constructs set-level diversity rewards via supergroup sampling and leave-one-out redistribution to expand the utility-diversity Pareto frontier in biomolecular design tasks.
-
Towards A Generative Protein Evolution Machine with DPLM-Evo
DPLM-Evo introduces an evolutionary discrete diffusion framework with explicit edit prediction and contextual noising that claims SOTA single-sequence mutation effect prediction on ProteinGym while supporting variable-length evolution simulation.
-
MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction
MKGR integrates region-aware sequence encoding with four protein-centered KGs via graph attention, bridge reconstruction, and pair gating to outperform baselines in cold-start PPI prediction.
-
Structure-Regularized Interpretable TCR-Epitope Prediction
TCR-SRIM uses structure regularization and contact prototypes for interpretable TCR-epitope binding prediction, reports SOTA performance on TCR-XAI, and finds generated structures produce less accurate interaction patterns than experimental ones.
-
Benchmarking virtual cell models for in-the-wild perturbation response
A new benchmarking framework shows virtual cell models overestimate performance on standard tests, drop sharply on unseen contexts and perturbations, and produce inconsistent rankings across metrics.
-
Polyformer: a generative framework for thermodynamic modeling of polymeric molecules
Polyformer generates sequence- and temperature-dependent conformational ensembles for proteins that agree with molecular dynamics simulations.
-
MolClaw: An Autonomous Agent with Hierarchical Skills for Drug Molecule Evaluation, Screening, and Optimization
MolClaw deploys a hierarchical skill architecture to reach state-of-the-art results on a new benchmark of multi-step drug discovery tasks.
-
SubQuad: Near-Quadratic-Free Structure Inference with Distribution-Balanced Objectives in Adaptive Receptor framework
SubQuad reports near-subquadratic immune-repertoire analysis with fairness-aware clustering, but its central performance and coverage claims are not supported by reproducible artifacts.