Pith. sign in

REVIEW 3 major objections 6 minor 87 references

Data-driven Discovery of Biophysical T Cell Receptor Co-specificity Rules

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that simple, interpretable distance metrics learned from T cell receptor sequence data predict whether two receptors share a peptide target, and that these rules transfer to peptides highly dissimilar from any seen in…

desk verdict A genuinely useful TCR distance metric with a real generalization claim, but the headline extrapolation rests on a subset whose size the paper never reports. read the letter →

arxiv 2412.13722 v3 pith:HSCUFWGY submitted 2024-12-18 q-bio.BM cs.LGphysics.bio-ph

classification q-bio.BMcs.LGphysics.bio-ph
keywords TcellreceptorTCRco-specificitycontrastivelearningdistancemetricaminoacidsubstitutionmatrixSARS-CoV-2two-pointstatistics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to show that the chance two T cell receptors bind the same peptide can be captured by a simple, interpretable distance metric learned from receptor sequence data, and that the resulting rules transfer to peptides very different from any seen during training. It derives the metric from a contrastive loss that it connects to pseudo-likelihood maximization, then fits position- and amino-acid-specific weights on a large set of SARS-CoV-2-specific TCRs. The learned joint metric outperforms the standard TCRdist metric on held-out and external data, and remains predictive for peptide targets differing by at least six amino acids from the nearest training peptide. The fitted weights indicate that matching the steric (shape) properties of substituted amino acids matters more for co-specificity than hydrophobic properties, and that residues not in direct peptide contact still contribute substantially. These findings support a view of TCR specificity as governed by local, transferable sequence rules.

What carries the argument

The machinery is a factorized, interpretable distance metric defined over edit steps between TCR CDR3 sequences. For each edit step $i$, the metric combines a site weight $S(k_i, \ell)$ and an amino-acid substitution weight $A(\Delta_i)$ through $d_{AS}(\sigma,\sigma') = -\sum_i \log\left[1 - \left(1 - e^{-S(k_i,\ell)}\right)\left(1 - e^{-A(\Delta_i)}\right)\right]$, which corresponds to a co-selection factor $Q_{AS}$ proportional to the product over sites of $[1 - P_S(k_i,\ell)(1 - P_A(\Delta_i))]$, with $P_S$ a per-site contribution probability and $P_A$ a per-substitution co-specificity retention probability. Parameters are fitted by minimizing a contrastive loss that the paper derives as a negative log-pseudo-likelihood, whose optimum sets each feature weight to the log-ratio of co-specific pair frequency to unlabeled pair frequency — a direct data-driven analog of log-odds substitution matrices. This two-point-statistics view is what lets the learned rules generalize across ligand landscapes rather than fitting one peptide at a time.

What would settle it

Measure co-specificity rates for TCR pairs that differ at two positions where each single substitution is known to preserve co-specificity: the factorized model (Eq. 14) predicts the double-substitution co-specificity should be the product of the two single-site probabilities; if the observed double-mutant co-specificity is substantially lower, the independence assumption fails and the metric's weights are mis-calibrated.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that TCR co-specificity — whether two receptors share a ligand — is governed by local and largely ligand-independent sequence rules that can be learned from data. Fitting an exponential model of co-selection factors, the authors optimize a factorized distance metric $d_{AS}$ that assigns separate weights to the position of an amino-acid substitution and the identity of the substitution, with the two combined as independent probabilities per edit step. This metric outperforms BLOSUM-based TCRdist on both held-out SARS-CoV-2 specificity groups and an external dataset covering other pathogens, and it continues to separate co-specific from cross-specific pairs when the peptide differs by at least six amino acids from any training peptide. The learned substitution matrix is only moderately correlated with BLOSUM62 and tracks steric properties more strongly than hydrophobicity, while the learned site weights show that positions not contacting the peptide matter for specificity. The paper concludes that simple learned distance metrics can substitute for heuristic ones and that contrastive learning provides a principled route to two-point statistics of receptor–ligand maps.

Load-bearing premise

Everything rests on the assumption that the impact of an amino-acid change depends only on the site and the amino-acid identity, with these two factors multiplying independently per edit step; if a substitution's effect changes with its sequence context, the learned weights are biased and the demonstrated extrapolation to multi-edit pairs is not assured.

Editorial extensions

If this is right

  • Learned TCR distance metrics can replace heuristic ones like TCRdist as drop-in similarity scores for clustering and annotation of TCR repertoires.
  • The same contrastive framework can be retrained on larger paired-chain data to yield a machine-optimized metric with more parameters and, presumably, further reach.
  • Because the learned substitution matrix emphasizes steric shape over hydrophobicity, sequence-similarity tools for T cells should be recalibrated accordingly rather than relying on BLOSUM.
  • The finding that non-contact CDR3 positions influence co-specificity argues against trimming loop edges when scoring TCR similarity.
  • Some co-specificity rules are shared between the $\alpha$ and $\beta$ chains, since metrics trained on CDR3$\beta$ transfer to CDR3$\alpha$ pairs in external data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence not explored in the paper: the same two-point, pseudo-likelihood machinery should transfer to antibody–antigen co-specificity data, where the learned substitution matrix could be compared directly.
  • Because the factorized model degrades past 3–4 edits, an explicit next step is to add pairwise, context-dependent terms to Eq. 14; if that restores performance at larger edit distances, the independence assumption is the bottleneck.
  • The paper's property analysis suggests a broader rule: flexible binding loops (TCR, antibody) are governed by steric matching, while rigid pockets (protease) are governed by hydrophobicity; this can be tested on additional protein–protein interaction classes.
  • The successful transfer from one virus's peptides to unrelated epitopes implies co-specificity rules could serve as a zero-shot prior for newly discovered epitopes with no known TCRs, something the paper mentions only as an eventual aim.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a supervised contrastive learning framework, derived from pseudo-likelihood maximization, to learn interpretable sequence distance metrics for TCR co-specificity. Using single-edit CDR3beta pairs from the MIRA SARS-CoV-2 dataset with matched V genes and restricted lengths, the authors learn an amino-acid substitution matrix A, a site-weight matrix S, and a joint factorized metric d_AS. They report that d_AS outperforms TCRdist on held-out MIRA pMHC groups and on external VDJdb data, including for TCR pairs with multiple edit steps. They also claim the learned rules generalize to pMHCs that differ by at least six amino acids from any training peptide. In addition, they find that steric properties of amino acids correlate more strongly with the learned substitution penalties than hydrophobic properties, and that non-contact CDR3beta positions contribute substantially to co-specificity.

Significance. If fully supported, this is a valuable contribution to TCR specificity prediction. The pseudo-likelihood derivation connects contrastive metric learning to established statistical physics inference, and the learned metrics are simple, interpretable, and available as drop-in replacements for TCRdist. The external validation on VDJdb and the transfer to TCRalpha chains strengthen the claim that local co-specificity rules are shared across ligands and chains. The authors provide public code and a clear data-availability statement, and they explicitly discuss several limitations, including the factorization assumption and the positive-unlabeled nature of the evaluation. The paper's central empirical assertion, however, requires stronger statistical support for the specific claim of generalization to pMHCs at least six amino acids away from the training set, which is currently not quantified at the group level.

major comments (3)
  1. [Section IV D, Fig. S5C] The headline empirical claim—that the learned rules generalize to pMHCs differing by at least six amino acids from the nearest training peptide—is not supported with sufficient quantitative detail. Fig. S5A shows a distribution of minimum peptide Levenshtein distances, but the paper does not report how many of the 24 held-out MIRA pMHC groups fall in the >=6 amino-acid category, nor how many TCR pairs contribute to that subset. The error bars in Fig. S5C are computed over resampled training/test batches, which resample within the same pMHC groups and therefore do not capture group-level variability. If only a few groups satisfy the >=6 condition, the reported AUROC could be driven by easy-to-discriminate groups, and the conclusion that the rules generalize broadly to dissimilar peptides would not be established. I request that the authors report the number of pMHC groups and TCR pairs in the dissimilar-peptide subset, and provide either per-group AUROC values or a bootstrap over pMHC groups (e.g., cluster bootstrap) to quantify group-level uncertainty.
  2. [Section III, Section VI] The evaluation treats all pairs from different specificity groups as negative, which is a positive-unlabeled (PU) setup because cross-reactive pairs are labeled as negative. The authors acknowledge this and note that about one in ten single edits maintain co-specificity, bounding AUROC below one. However, this labeling issue can also affect the comparison between metrics if the rate of cross-reactivity differs across edit distances or between the MIRA and VDJdb test sets. For instance, at edit distance 1 the mislabeling rate may be higher than at distance 5, and if the composition of true positives differs between datasets, the reported AUROC gap between d_AS and TCRdist could be partly an artifact of label noise rather than metric quality. I recommend a sensitivity analysis that either restricts to high-confidence negatives (e.g., pairs with large TCR sequence dissimilarity or different HLA restrictions) or estimates and corrects for mislabeling rates, to confirm that the relative improvements are robust.
  3. [Section IV D, Eq. 14-15] The joint metric d_AS assumes that site contributions and amino-acid substitution effects factorize, i.e., that P_S and P_A are statistically independent. The authors mention this assumption in the Discussion, but the empirical section does not test it. Because the metric's extrapolation to multi-edit pairs and to dissimilar pMHCs relies on this factorization, I ask the authors to report how the model's performance compares with a non-factorized alternative (e.g., a full site-by-amino-acid matrix trained on the same data) or to provide diagnostic evidence that the independence assumption is not the main cause of the observed performance decline at edit distances beyond 3-4. This would strengthen the interpretation of the learned weights as biophysical quantities rather than as phenomenological fit parameters.
minor comments (6)
  1. [Section I] In the sentence 'P(σ) represents the highly highly non-uniform', 'highly' is repeated; this typo should be corrected.
  2. [Figure S3 and S4 captions] The captions state 'resampling training data with replacement fiveteen times'; 'fiveteen' should be 'fifteen'.
  3. [Section V] The word 'biopysical' in 'biopysical interpretation' should be 'biophysical'.
  4. [Table S1] In the Steric Property rows, 'STERMIMOL' should be 'STERIMOL' (three occurrences).
  5. [Section IV D, Eq. 14] The notation switches between A(Δ_i) in Eq. 12 and P_A(Δ_i) in Eq. 15; please define the relationship explicitly to avoid confusion when interpreting the learned matrix A.
  6. [Figure S5] Fig. S5C reports error bars 'across 10 test batches', while other figures use 15 resampled batches; please clarify why the number differs and ensure the caption describes the resampling procedure consistently.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: learned metrics are fit on training pMHC groups and evaluated on held-out and external data, so the generalization claim is out-of-sample.

full rationale

The paper's central derivation is empirical risk minimization: parameters of d_A, d_S, and d_AS are optimized by minimizing the contrastive/pseudo-likelihood loss (Eq. 7) on positive pairs from 26 SARS-CoV-2 MIRA pMHC groups and unlabeled pairs. The headline generalization claims are then evaluated on 24 MIRA pMHC groups excluded from training and on VDJdb groups with pMHCs not present in MIRA, with AUROCs computed over resampled balanced batches. This is a genuine out-of-sample evaluation, not a refit of test labels. The factorized ansatz in Eqs. 14-15 is explicitly presented as a modeling assumption ('Most naively', 'Assuming further ... statistically independent'), and its parameters are freely fitted, so the reported predictive performance is not forced by construction. Self-citations (refs. 20, 26-28) supply empirical context, such as the previously observed Levenshtein-distance decay of co-selection factors, but these are external, falsifiable results rather than uniqueness theorems or fitted parameters imported into the loss; the optimized matrices and their held-out performance do not reduce to those citations. The structural contact-probability comparison (r = 0.47) and amino-acid property correlations are independent external analyses rather than outputs of the optimization. Concerns about the possibly small subset of test pMHCs at >=6 amino-acid distance (Fig. S5) are statistical robustness issues, not definitional circularity, and do not change the circularity score.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim rests on a fitted substitution matrix and fitted site weights, on an exponential ansatz and independence factorization for the co-selection probability, and on a data curation scheme that restricts the training regime. No new physical entities are postulated.

free parameters (3)
  • Amino-acid substitution matrix A (with gap) / P_A(Delta) = 20x20 plus gap, learned via gradient descent; values shown in Fig. 4A and Fig. S1C
    Central fitted object: assigns co-specificity penalties to each amino-acid substitution; used in d_A and d_AS and drives the steric-vs-hydrophobic interpretation.
  • Site weight matrix S(k,l) = Length-dependent position weights for CDR3beta lengths 11-15; shown in Fig. 3A and Fig. S2
    Central fitted object: assigns co-specificity penalties to substitution position; used in d_S and d_AS and drives the non-contact-position interpretation.
  • Training hyperparameters (learning rate, steps, batch resampling) = lr=1e-3; 3000-12000 Adam steps; 15 subsampled batches of 1000 per pMHC
    Chosen by hand, not fitted to data, but they affect the inferred matrices and should be fixed for reproducibility.
assumptions (5)
  • domain assumption Exponential ansatz: average co-selection factor Q(sigma,sigma') is proportional to e^{-d_theta(sigma,sigma')} (Eq. 4).
    Assumes the probability that two TCRs share specificity decays exponentially with the learned distance; the functional form is not derived from a binding model, and all subsequent parameter fitting depends on it.
  • domain assumption Ligand averaging: Q_{pi_i}(sigma)Q_{pi_i}(sigma') approximately equals Q(sigma,sigma'), the average over ligands (Eq. 3).
    Assumes co-selection factors are sufficiently universal across pMHCs that ligand-specific variation can be ignored when learning general rules; the authors acknowledge ligand-dependent variation exists and deliberately average over it.
  • domain assumption Factorization of substitution contributions: Q_AS(sigma,sigma') is proportional to the product over edit steps of [1 - P_S(k_i,l)(1-P_A(Delta_i))] (Eq. 15).
    Assumes each edit step contributes independently and site- and identity-dependent probabilities are statistically independent; the main structural risk for the metric's extrapolation to multi-edit pairs.
  • domain assumption Empirical unlabeled-set average approximates the generative pair distribution: the denominator in Eq. 5-7 is computed from the same training pairs, including co-specific pairs as a regularizer.
    The normalization depends on the training composition; the paper treats this as regularizing based on prior work (refs. 20, 38).
  • ad hoc to paper Data curation exclusions (CDR3beta lengths 11-15, identical V genes, no cysteine, single-edit training pairs) preserve the co-specificity signal of interest (Section III).
    These restrictions reduce confounders but also narrow the regime; no analysis shows the learned rules transfer to excluded sequence classes such as different lengths or V-gene mismatches.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Data-driven Discovery of Biophysical T Cell Receptor Co-specificity Rules." pith.science (2026). https://pith.science/paper/HSCUFWGY

@misc{pith2026241213722,
  author       = {Pith},
  title        = {Pith review of: Data-driven Discovery of Biophysical T Cell Receptor Co-specificity Rules},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HSCUFWGY}},
  note         = {Machine review of arXiv:2412.13722}
}
read the original abstract

The biophysical interactions between the T cell receptor (TCR) and its ligands determine the specificity of the cellular immune response. However, the immense diversity of receptors and ligands has made it challenging to discover generalizable rules across the distinct binding affinity landscapes created by different ligands. Here, we present an optimization framework for discovering biophysical rules that predict whether TCRs share specificity to a ligand. Applying this framework to TCRs associated with a collection of SARS-CoV-2 peptides we systematically characterize how co-specificity depends on the type and position of amino-acid differences between receptors. We also demonstrate that the inferred rules generalize to ligands highly dissimilar to any seen during training. Our analysis reveals that matching of steric properties between substituted amino acids is more important for receptor co-specificity than the hydrophobic properties that prominently determine evolutionary substitutability. Our analysis also quantifies the substantial importance of positions not in direct contact with the peptide for specificity. These findings highlight the potential for data-driven approaches to uncover the molecular mechanisms underpinning the specificity of adaptive immune responses.

Figures

Figures reproduced from arXiv: 2412.13722 by the authors.

Figure 1
Figure 1. Contrastive learning of rules that generalize across complex receptor-ligand maps. A) Cartoon disordered “landscape” of selection factors Qπ(σ), which describe the varying binding affinities of TCRs σ to ligands π. Averaging selection factors over ligands, ⟨Qπ(σ)⟩P (π), will lead to a largely flat marginal distribution (inset), unless we are able to restrict the average to ligands with highly similar TCR landscapes.… view at source ↗
Figure 2
Figure 2. Learning of a CDR3β distance metric that generalizes to unseen ligands. A) Distance metrics are defined in terms of the edit steps between two CDR3β sequences. Each edit step includes the length of sequence ℓ, position of substitution κ, and the identity of substitution ∆. B) Receiver Operating Characterics (ROC) curves for identifying co-specific relative to cross-specific (i.e., ∈ C = {(σi, σj )|πi ̸= πj , ∀i, j})… view at source ↗
Figure 3
Figure 3. Comparison of site-dependent weights with TCR-pMHC contact probabilities reveals the importance of non-contact sites. A) Site weights, PS, learned from learned from joint optimization of substitution type-dependent and site-dependent weights (dAS). Sites that lack sufficient substitution statistics are shown in gray. B) Average contact probability in TCR-pMHC crystal structures (see Methods). Sites along the CDR3β s… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Optimized amino-acid substitution matrix differs from BLOSUM. A) Probability of maintaining co￾specificity, PA, learned from joint optimization of substitution type-dependent and site-dependent weights (dAS). Note that cysteine was removed from the optimization due to …
Figure 5
Figure 5. Figure 5: Physical correlates of the optimized amino-acid substitution matrix highlight the contribution of shape to TCR specificity. A) Spearman’s rank correlation coefficient (r) of substitution matrices with matrices constructed from pairwise absolute difference in physical p…
Figure 6
Figure 6. Figure 6: Learned metrics generalize to pairs of TCRs with a range of sequence similarities. AUROC of baseline method TCRdist (dTd, black) and inferred site-specific metric (dS, green), amino-acid substitution metric (dA, blue), and jointly learned metric (dAS, red) for edit dis…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

87 extracted references · 76 canonical work pages

  1. [1]

    M. M. Davis and P. J. Bjorkman, The T cell receptor genes and T cell recognition, Nature334, 395 (1988)

  2. [2]

    H. Chi, M. Pepper, and P. G. Thomas, Principles and therapeutic applications of adaptive immunity, Cell187, 2052 (2024)

  3. [3]

    P. Dash, A. J. Fiore-Gartland, T. Hertz, G. C. Wang, S. Sharma, A. Souquette, J. C. Crawford, E. B. Clemens, T. H. O. Nguyen, K. Kedzierska, N. L. La Gruta, P. Bradley, and P. G. Thomas, Quantifiable predictive features define epitope-specific T cell receptor reper- toires, Nature547, 89 (2017)

  4. [4]

    Glanville, H

    J. Glanville, H. Huang, A. Nau, O. Hatton, L. E. Wa- gar, F. Rubelt, X. Ji, A. Han, S. M. Krams, C. Pettus, N. Haas, C. S. L. Arlehamn, A. Sette, S. D. Boyd, T. J. Scriba, O. M. Martinez, and M. M. Davis, Identifying specificity groups in the T cell receptor repertoire, Na- ture547, 94 (2017)

  5. [5]

    C. S. Dobson, A. N. Reich, S. Gaglione, B. E. Smith, E. J. Kim, J. Dong, L. Ronsard, V. Okonkwo, D. Ling- wood, M. Dougan, S. K. Dougan, and M. E. Birnbaum, Antigen identification and high-throughput interaction mapping by reprogramming viral entry, Nature Methods 19, 449–460 (2022)

  6. [6]

    A. V. Joglekar, M. T. Leonard, J. D. Jeppson, M. Swift, G. Li, S. Wong, S. Peng, J. M. Zaretsky, J. R. Heath, A. Ribas,et al., T cell antigen discovery via signaling and antigen-presenting bifunctional receptors, Nature meth- ods16, 191 (2019)

  7. [7]

    Sureshchandra, J

    S. Sureshchandra, J. Henderson, E. Levendosky, S. Bhat- tacharyya, J. M. Kastenschmidt, A. M. Sorn, M. T. Mi- tul, A. Benchorin, K. Batucal, A. Daugherty,et al., Tis- sue determinants of the human T cell receptor repertoire, bioRxiv preprint 10.1101/2024.08.17.608295 (2024)

  8. [8]

    D. V. Bagaev, R. M. A. Vroomans, J. Samir, U. Stervbo, C. Rius, G. Dolton, A. Greenshields-Watson, M. Attaf, E. S. Egorov, I. V. Zvyagin, N. Babel, D. K. Cole, A. J. Godkin, A. K. Sewell, C. Kesmir, D. M. Chudakov, F. Lu- ciani, and M. Shugay, VDJdb in 2019: database exten- sion, new analysis infrastructure and a T-cell receptor motif compendium, Nucleic ...

Show all 87 references
  1. [9]

    Tickotsky, T

    N. Tickotsky, T. Sagiv, J. Prilusky, E. Shifrut, and N. Friedman, McPAS-TCR: a manually curated cata- logue of pathology-associated T cell receptor sequences, Bioinformatics33, 2924–2929 (2017)

  2. [10]

    R. Vita, S. Mahajan, J. A. Overton, S. K. Dhanda, S. Martini, J. R. Cantrell, D. K. Wheeler, A. Sette, and B. Peters, The immune epitope database (IEDB): 2018 update, Nucleic Acids Research47, D339–D343 (2018)

  3. [11]

    Nolan, M

    S. Nolan, M. Vignali, M. Klinger, J. N. Dines, I. M. Ka- plan, E. Svejnoha, T. Craft, K. Boland, M. W. Pesesky, R. M. Gittelman,et al., A large-scale database of T-cell receptor beta sequences and binding associations from natural and synthetic exposure to SARS-CoV-2, Fron- ti...

  4. [12]

    Gielis, P

    S. Gielis, P. Moris, W. Bittremieux, N. De Neuter, B. Ogunjimi, K. Laukens, and P. Meysman, Detection of Enriched T Cell Epitope Specificity in Full T Cell Receptor Sequence Repertoires, Frontiers in Immunology 10(2019)

  5. [13]

    X. Lin, J. T. George, N. P. Schafer, K. Ng Chau, M. E. Birnbaum, C. Clementi, J. N. Onuchic, and H. Levine, Rapid assessment of T-cell receptor specificity of the im- mune repertoire, Nature Computational Science1, 362 (2021)

  6. [14]

    D. S. Fischer, Y. Wu, B. Schubert, and F. J. Theis, Predicting antigen specificity of single T cells based on TCR CDR3 regions, Molecular Systems Biology16, e9416 (2020)

  7. [15]

    Jokinen, J

    E. Jokinen, J. Huuhtanen, S. Mustjoki, M. Heinonen, and H. L¨ ahdesm¨ aki, Predicting recognition between T cell receptors and epitopes with TCRGP, PLOS Compu- tational Biology17, e1008814 (2021)

  8. [16]

    Croce, S

    G. Croce, S. Bobisse, D. L. Moreno, J. Schmidt, P. Guil- lame, A. Harari, and D. Gfeller, Deep learning predic- tions of TCR-epitope interactions reveal epitope-specific chains in dual alpha T cells, Nature Communications15, 10.1038/s41467-024-47461-8 (2024)

  9. [17]

    A. Wang, X. Lin, K. N. Chau, J. N. Onuchic, H. Levine, and J. T. George, RACER-m leverages structural fea- tures for sparse T cell specificity prediction, Science Ad- vances10, 10.1126/sciadv.adl0161 (2024)

  10. [18]

    Meynard-Piganeau, C

    B. Meynard-Piganeau, C. Feinauer, M. Weigt, A. M. Walczak, and T. Mora, TULIP: a transformer-based un- supervised language model for interacting peptides and T cell receptors that generalizes to unseen epitopes, Proceedings of the National Academy of Sciences121, e2316401121 (2024)

  11. [19]

    Zhang, Z

    Y. Zhang, Z. Wang, Y. Jiang, D. R. Littler, M. Ger- stein, A. W. Purcell, J. Rossjohn, H.-Y. Ou, and J. Song, Epitope-anchored contrastive transfer learning for paired CD8+ T cell receptor–antigen recognition, Nature Ma- chine Intelligence6, 1344–1358 (2024)

  12. [20]

    Nagano, A

    Y. Nagano, A. G. Pyo, M. Milighetti, J. Henderson, J. Shawe-Taylor, B. Chain, and A. Tiffeau-Mayer, Con- trastive learning of T cell receptor representations, Cell 12 Systems16(2025)

  13. [21]

    Nielsen, A

    M. Nielsen, A. Eugster, M. F. Jensen, M. Goel, A. Tiffeau-Mayer, A. Pelissier, S. Valkiers, M. R. Mart ´ ınez, B. Meynard-Piganeeau, V. Greiff, T. Mora, A. M. Walczak, G. Croce, D. L. Moreno, D. Gfeller, P. Meysman, and J. Barton, Lessons learned from the IMMREP23 TCR-epitope ...

  14. [22]

    Bellet, A

    A. Bellet, A. Habrard, and M. Sebban, A survey on metric learning for feature vectors and structured data (2014), arXiv:1306.6709 [cs.LG]

  15. [23]

    Vinyals, C

    O. Vinyals, C. Blundell, T. Lillicrap, K. Kavukcuoglu, and D. Wierstra, Matching networks for one shot learning (2016)

  16. [24]

    Henikoff and J

    S. Henikoff and J. G. Henikoff, Amino acid substitution matrices from protein blocks., Proceedings of the Na- tional Academy of Sciences of the United States of Amer- ica89, 10915 (1992)

  17. [25]

    Milighetti, J

    M. Milighetti, J. Shawe-Taylor, and B. Chain, Predict- ing T cell receptor antigen specificity from structural fea- tures derived from homology models of receptor-peptide- major histocompatibility complexes, Frontiers in Physi- ology12, 730908 (2021)

  18. [26]

    Mayer and C

    A. Mayer and C. G. Callan, Measures of epitope binding degeneracy from T cell receptor repertoires, Proceedings of the National Academy of Sciences120, e2213264120 (2023)

  19. [27]

    Tiffeau-Mayer, Unbiased estimation of sampling vari- ance for Simpson’s diversity index, Physical Review E 109, 064411 (2024)

    A. Tiffeau-Mayer, Unbiased estimation of sampling vari- ance for Simpson’s diversity index, Physical Review E 109, 064411 (2024)

  20. [28]

    Henderson, Y

    J. Henderson, Y. Nagano, M. Milighetti, and A. Tiffeau- Mayer, Limits on inferring T cell specificity from partial information, Proceedings of the National Academy of Sci- ences121, e2408696121 (2024)

  21. [29]

    Drost, L

    F. Drost, L. Schiefelbein, and B. Schubert, meTCRs - Learning a metric for T-cell receptors, bioRxiv preprint , 2022.10.24.513533 (2022)

  22. [30]

    Pertseva, O

    M. Pertseva, O. Follonier, D. Scarcella, and S. T. Reddy, TCR clustering by contrastive learning on antigen speci- ficity, Briefings in Bioinformatics25, bbae375 (2024)

  23. [31]

    Murugan, T

    A. Murugan, T. Mora, A. M. Walczak, and C. G. Callan, Statistical inference of the generation probability of T- cell receptors from sequence repertoires, Proceedings of the National Academy of Sciences109, 16161 (2012)

  24. [32]

    Sethna, G

    Z. Sethna, G. Isacchini, T. Dupic, T. Mora, A. M. Wal- czak, and Y. Elhanati, Population variability in the gen- eration and selection of T-cell repertoires, PLOS Com- putational Biology16, e1008394 (2020)

  25. [33]

    Elhanati, A

    Y. Elhanati, A. Murugan, C. G. Callan, T. Mora, and A. M. Walczak, Quantifying Selection in Immune Recep- tor Repertoires, Proceedings of the National Academy of Sciences111, 9875 (2014)

  26. [34]

    Bravi, A

    B. Bravi, A. Di Gioacchino, J. Fernandez-de Cossio-Diaz, A. M. Walczak, T. Mora, S. Cocco, and R. Monasson, A transfer-learning approach to predict antigen immuno- genicity and T-cell receptor specificity, ELife12, e85126 (2023)

  27. [35]

    S. F. Edwards and P. W. Anderson, Theory of spin glasses, Journal of Physics F: Metal Physics5, 965 (1975)

  28. [36]

    Khosla, P

    P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan, Supervised Contrastive Learning, arXiv 10.48550/arXiv.2004.11362 (2021)

  29. [37]

    Ekeberg, C

    M. Ekeberg, C. L¨ ovkvist, Y. Lan, M. Weigt, and E. Au- rell, Improved contact prediction in proteins: using pseu- dolikelihoods to infer Potts models, Physical Review E 87, 012707 (2013)

  30. [38]

    Wang and P

    T. Wang and P. Isola, Understanding Contrastive Repre- sentation Learning through Alignment and Uniformity on the Hypersphere, arXiv 10.48550/arXiv.2005.10242 (2022)

  31. [39]

    S. F. Altschul, Amino acid substitution matrices from an information theoretic perspective, Journal of molecular biology219, 555 (1991)

  32. [40]

    Bekker and J

    J. Bekker and J. Davis, Learning from positive and unla- beled data: A survey, Machine Learning109, 719 (2020)

  33. [41]

    Goncharov, D

    M. Goncharov, D. Bagaev, D. Shcherbinin, I. Zvyagin, D. Bolotin, P. G. Thomas, A. A. Minervina, M. V. Pogorelyy, K. Ladell, J. E. McLaren, D. A. Price, T. H. O. Nguyen, L. C. Rowntree, E. B. Clemens, K. Kedzierska, G. Dolton, C. R. Rius, A. Sewell, J. Samir, F. Luciani, K. V. ...

  34. [42]

    W. D. Chronister, A. Crinklaw, S. Mahajan, R. Vita, Z. Ko¸ salo˘ glu-Yal¸ cın, Z. Yan, J. A. Greenbaum, L. E. Jessen, M. Nielsen, S. Christley, L. G. Cowell, A. Sette, and B. Peters, TCRMatch: predicting T-cell receptor specificity based on sequence similarity to previously ch...

  35. [43]

    Mayer-Blackwell, S

    K. Mayer-Blackwell, S. Schattgen, L. Cohen-Lavi, J. C. Crawford, A. Souquette, J. A. Gaevert, T. Hertz, P. G. Thomas, P. Bradley, and A. Fiore-Gartland, TCR meta- clonotypes for biomarker discovery with tcrdist3 enabled identification of public, HLA-restricted clusters of SARS...

  36. [44]

    V. I. Levenshtein, Binary Codes Capable of Correcting Deletions, Insertions and Reversals, Soviet Physics Dok- lady10, 707 (1966)

  37. [45]

    Koˇ smrlj, A

    A. Koˇ smrlj, A. K. Jha, E. S. Huseby, M. Kardar, and A. K. Chakraborty, How the thymus designs antigen- specific and self-tolerant T cell receptor sequences, Pro- ceedings of the National Academy of Sciences105, 16671 (2008)

  38. [46]

    J. Leem, S. H. P. de Oliveira, K. Krawczyk, and C. M. Deane, STCRDab: the structural T-cell receptor database, Nucleic Acids Research46, D406–D412 (2017)

  39. [47]

    H. Mei, Z. H. Liao, Y. Zhou, and S. Z. Li, A new set of amino acid descriptors and its application in peptide QSARs, Peptide Science80, 775–786 (2005)

  40. [48]

    Jankauskait˙ e, B

    J. Jankauskait˙ e, B. Jim´ enez-Garc ´ ıa, J. Dapk¯ unas, J. Fern´ andez-Recio, and I. H. Moal, SKEMPI 2.0: an updated benchmark of changes in protein–protein bind- ing energy, kinetics and thermodynamics upon mutation, Bioinformatics35, 462–469 (2018)

  41. [49]

    Postovskaya, K

    A. Postovskaya, K. Vercauteren, P. Meysman, and K. Laukens, tcrBLOSUM: an amino acid substitution matrix for sensitive alignment of distant epitope-specific TCRs, Briefings in Bioinformatics26, bbae602 (2025)

  42. [50]

    A. S. Perelson and G. F. Oster, Theoretical studies of clonal selection: minimal antibody repertoire size and reliability of self-non-self discrimination, Journal of the- oretical biology81, 645 (1979). 13

  43. [51]

    B. D. Stadinski, K. Shekhar, I. G´ omez-Touri˜ no, J. Jung, K. Sasaki, A. K. Sewell, M. Peakman, A. K. Chakraborty, and E. S. Huseby, Hydrophobic CDR3 residues promote the development of self-reactive T cells, Nature immunol- ogy17, 946 (2016)

  44. [52]

    V. K. Karnaukhov, D. S. Shcherbinin, A. O. Chugunov, D. M. Chudakov, R. G. Efremov, I. V. Zvyagin, and M. Shugay, Structure-based prediction of T cell recep- tor recognition of unseen epitopes using TCRen, Nature Computational Science4, 510 (2024)

  45. [53]

    A. Y. Leary, D. Scott, N. T. Gupta, J. C. Waite, D. Skokos, G. S. Atwal, and P. G. Hawkins, Designing meaningful continuous representations of T cell receptor sequences with deep generative models, Nature Commu- nications15, 10.1038/s41467-024-48198-0 (2024)

  46. [54]

    H. F. Jones, Z. Molvi, M. G. Klatt, T. Dao, and D. A. Scheinberg, Empirical and rational design of T cell receptor-based immunotherapies, Frontiers in Immunol- ogy11, 10.3389/fimmu.2020.585385 (2021)

  47. [55]

    Paszke, S

    A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, Automatic differentiation in PyTorch, Conference on Neural Information Processing System (2017)

  48. [56]

    Pyo, Inference code for TCR co-specificity prediction, Zenodo repository 10.5281/zenodo.15707892 (2025)

    A. Pyo, Inference code for TCR co-specificity prediction, Zenodo repository 10.5281/zenodo.15707892 (2025). 14 SUPPLEMENT AR Y FIGURES AND T ABLES Figure S1.Comparison of amino-acid substitution matrices (supplement to Fig. 4). A)Substitution penalties, ˜dTd, used in TCRdist.B...

  49. [57]

    Free energy of solution in water

  50. [58]

    Solvation free energy

  51. [59]

    Number of hydrogen-bond donors

  52. [60]

    Number of full nonbonding orbitals

  53. [61]

    Retention coefficient in high performance liquid chromatography (HPLC), pH 7.4

  54. [62]

    Retention coefficient in HPLC, pH 2.1

  55. [63]

    Partition coefficient in thin-layer chromatography

  56. [64]

    Retention coefficient at pH 2 13.R f for 1-N-(4-nitrobenzofurazono)-amino acids in ethyl acetate/pyridine/water

  57. [65]

    ∆Gof transfer from organic solvent to water

  58. [66]

    Hydration potential or free energy of transfer from vapor phase to water 16.R f salt chromatography

  59. [67]

    logD, partition coefficient at pH 7.1 for acetylamide derivatives of amino acids in octanol/water

  60. [68]

    Average volume of buried residue

    ∆G=RTlogf,f= fraction buried/accessible amino acids in 22 proteins Steric Property 19. Average volume of buried residue

  61. [69]

    Residue accessible surface area in tripeptide

  62. [70]

    Normalized van der Waals volume

  63. [71]

    STERMIMOL length of the side chain

  64. [72]

    STERMIMOL minimum width of the side chain

  65. [73]

    STERMIMOL maximum width of the side chain

  66. [74]

    Average accessible surface area

  67. [75]

    Distance between C α and centroid of side chain

  68. [76]

    Side-chain torsion angleϕ

  69. [77]

    Radius of gyration of side chain

  70. [78]

    Van der Waals parameterR 0

  71. [79]

    Van der Waals parameterϵ

  72. [80]

    Substituent van der Waals volume Electronic Property 36.αCH chemical shifts 37.αNH chemical shifts

  73. [81]

    A parameter of charge transfer capability

  74. [82]

    A parameter of charge transfer donor capability

  75. [83]

    Nuclear magnetic resonance (NMR) chemical shift ofαcarbon

  76. [84]

    Localized electrical effect

  77. [85]

    Amphipathicity index

  78. [86]

    Electron-ion interaction potential values

  79. [87]

    Table of amino-acid properties

    pKCOOH(COOH on C α) Table SI. Table of amino-acid properties. This table lists the physical and chemical properties of amino acids used for comparison with substitution matrices in Fig. 5. The values were obtained from Mei et al.,Biopolymers(2005)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.