Pith. sign in

REVIEW 14 cited by

Deciphering antibody affinity maturation with language models and weakly supervised learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.07782 v1 pith:MRZVQWVZ submitted 2021-12-14 q-bio.BM cs.LG

classification q-bio.BMcs.LG
keywords antibodiesimmunelanguagemodelssequencesaffinityantibodybinding
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In response to pathogens, the adaptive immune system generates specific antibodies that bind and neutralize foreign antigens. Understanding the composition of an individual's immune repertoire can provide insights into this process and reveal potential therapeutic antibodies. In this work, we explore the application of antibody-specific language models to aid understanding of immune repertoires. We introduce AntiBERTy, a language model trained on 558M natural antibody sequences. We find that within repertoires, our model clusters antibodies into trajectories resembling affinity maturation. Importantly, we show that models trained to predict highly redundant sequences under a multiple instance learning framework identify key binding residues in the process. With further development, the methods presented here will provide new insights into antigen binding from repertoire sequences alone.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AbRank: A Benchmark Dataset and Metric-Learning Framework for Antibody-Antigen Affinity Ranking

    q-bio.BM 2025-06 conditional novelty 6.0 of 10

    AbRank aggregates over 380,000 binding assays into pairwise ranking benchmarks with graded generalization splits, and shows that a ranking-trained graph network, WALLE-Affinity, often beats regression-trained and stru...

  2. NbBench: Benchmarking Language Models for Comprehensive Nanobody Tasks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    NbBench provides the first unified nanobody benchmark and shows that antibody language models lead on binding tasks while affinity and stability regressions remain hard for all tested models.

  3. Steering Protein Family Design through Profile Bayesian Flow

    q-bio.BM 2025-02 conditional novelty 6.0 of 10

    ProfileBFN adapts Bayesian flow networks to accept protein-family profiles, enabling diverse, novel, and apparently functional family protein generation from single-sequence training.

  4. Energy-based generative models for monoclonal antibodies

    q-bio.BM 2024-11 conditional novelty 6.0 of 10

    Energy-based sampling with a human-antibody prior generates diverse heavy-chain mutants near the wild type that lie on predicted affinity-solubility Pareto fronts and outperform constrained local search in synthetic b...

  5. Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design

    q-bio.BM 2026-07 reject novelty 5.0 of 10

    AAMFM combines ESM3, an antigen-geometry adapter, and Cal-DPO preference optimization rewarded by AlphaFold3-style scores to design antibody CDRs and structures, reporting higher predicted binding scores than prior methods.

  6. Conditionally Site-Independent Neural Evolution of Antibody Sequences

    cs.LG 2026-02 conditional novelty 5.0 of 10

    A neural continuous-time Markov model of antibody affinity maturation that beats language models on fitness prediction and steers sampling toward antigen-specific binders.

  7. Antibody Design and Optimization with Multi-scale Equivariant Graph Diffusion Models for Accurate Complex Antigen Binding

    cs.LG 2025-06 conditional novelty 5.0 of 10

    AbMEGD, a fusion of ViS-MP and IPA inside a diffusion process, reports modest CDR-H3 gains over DiffAb on SAbDab.

  8. AffinityFlow: Guided Flows for Antibody Affinity Maturation

    cs.LG 2025-02 reject novelty 5.0 of 10

    AffinityFlow guides AlphaFlow structure generation toward low Rosetta binding energy, then inverse-folds the structures to propose antibody mutations, and reports top scores on a computational affinity maturation benchmark.

  9. Harnessing Preference Optimisation in Protein LMs for Hit Maturation in Cell Therapy

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A preference-fine-tuned protein language model ranks single and double mutants of CAR binding domains better than the pretrained model, finding several mutants that outperform their parent.

  10. S$^2$ALM: Sequence-Structure Pre-trained Large Language Model for Comprehensive Antibody Representation Learning

    cs.LG 2024-11 conditional novelty 5.0 of 10

    S2ALM pretrains a 650M-parameter antibody language model on sequence and Foldseek 3Di structure tokens, reporting modest state-of-the-art gains across six antibody benchmarks.

  11. Llama-Affinity: A Predictive Antibody Antigen Binding Model Integrating Antibody Sequences with Llama3 Backbone Architecture

    cs.LG 2025-05 conditional novelty 3.0 of 10

    A Llama-style transformer trained on antibody sequences is reported to classify antibody-antigen binding with 0.964 accuracy and 0.994 ROC AUC, outperforming published baselines.

  12. Computational Protein Science in the Era of Large Language Models (LLMs)

    cs.CE 2025-01 conditional novelty 3.0 of 10

    A survey that categorizes protein language models by the knowledge they learn and reviews their applications, with no new experimental results.

  13. A Comprehensive Review of Protein Language Models

    q-bio.BM 2025-02 conditional novelty 2.0 of 10

    A survey paper that catalogs protein language models, their architectures, training data, benchmarks, and tools, but lacks a systematic methodology and contains several factual errors.

  14. Sequence-based protein-protein interaction prediction and its applications in drug discovery

    q-bio.BM 2025-07 conditional

    A comprehensive survey of sequence-based protein-protein interaction prediction methods, their evaluation pitfalls, and their applications in drug discovery.

Pith tools